Scheduling method for HCPS workshop system based on multi-agent deep reinforcement learning
Through the multi-agent deep reinforcement learning method, a worker efficiency fluctuation model and an ARSI training mechanism are built, which solves the problem of insufficient consideration of worker factors by single-agent algorithm, achieves more efficient production scheduling and optimization of worker factors, and improves the accuracy and stability of the production process.
Patent Information
- Application Number
- CN202510387466.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-08
AI Technical Summary
The existing single agent reinforcement learning algorithm has limitations in taking into account worker factors, and it is difficult to effectively deal with complex production scheduling needs, especially the differences in workers' fatigue and skills.
The multi-agent deep reinforcement learning method is adopted to build a worker efficiency fluctuation model, design a Markov decision-making process, introduce an ARSI training mechanism, and optimize production scheduling through the collaboration between workpiece agents and worker agents.
The accuracy and efficiency of workshop scheduling are significantly improved, especially by modeling each worker individually to consider their fatigue and skill factors, optimizing the production process, and improving the collaboration efficiency and system stability between the workpiece agent and the worker agent.
Smart Images

Figure CN120278459A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of production scheduling, and particularly relates to a scheduling method for an HCPS workshop system based on multi-agent deep reinforcement learning. Background Art
[0002] HCPS (Human-Cyber-Psysical-system, human, information and physical system) emphasizes the deep integration of humans, machines and information, and this concept is of great significance in intelligent manufacturing. In the intelligent manufacturing environment, humans not only play a leading role in the decision-making process, but also need to reasonably plan and allocate the workload to ensure the balance between production efficiency and personnel burden. Therefore, the production mode combining human-machine collaboration has become the core trend of intelligent manufacturing.
[0003] In the current production scheduling planning, reinforcement learning algorithms such as DQN, PPO, DDPG, etc. are widely used to solve the workshop scheduling problem. These algorithms usually belong to single-agent algorithms and can effectively handle some simple production scheduling problems. However, with the increasing complexity of production requirements, more and more producers hope to consider the factors of workers more comprehensively in the scheduling planning. For example, considering the change of workers' fatigue degree and the differential operation ability between workers, the capabilities of single-agent algorithms show limitations and are difficult to fully meet the changing and complex production scheduling requirements. Therefore, further optimizing the scheduling algorithm, especially starting from the factors of workers, has become an important direction to improve production efficiency and ensure the working environment. Summary of the Invention
[0004] The present invention aims to overcome the problem of insufficient consideration of worker factors in the existing scheduling scheme, and proposes a scheduling method for an HCPS workshop system based on multi-agent deep reinforcement learning.
[0005] To solve the above technical problems, the present invention is realized through the following technical solutions:
[0006] A scheduling method for an HCPS workshop system based on multi-agent deep reinforcement learning includes the following steps:
[0007] Step 1, construct an objective optimization model with the goal of minimizing the completion time;
[0008] Step 2, establish a worker efficiency fluctuation model for each worker separately;
[0009] Step 3, model the HCPS workshop system as a Markov decision process and design decision points;
[0010] Step 4, design the state space;
[0011] Step 5, design the action space:
[0012] Step 6, design a reward function based on the ARSI training mechanism.
[0013] Furthermore, the objective optimization model formula in Step 1 is as follows:
[0014]
[0015] Among them, C max represents the total completion time of all processing tasks, x represents the workpiece, n is the total number of workpieces, and ETx is the completion time of each workpiece.
[0016] Furthermore, in Step 2, establish a worker efficiency fluctuation model, which specifically includes:
[0017] Step 21, initialize the worker fatigue accumulation rate, fatigue recovery rate, skill level of each worker, as well as the working state and fatigue degree at each stage;
[0018] Step 22, consider the fatigue accumulation of workers. The change in fatigue degree considers the fatigue accumulation caused by workers' work. The fatigue accumulation calculation formula is as follows:
[0019]
[0020] Among them, f w (t i ) is the fatigue degree of worker w at the current moment; λ is the fatigue accumulation rate; t i represents the current moment; t i-1 represents the previous moment; i represents the index of time;
[0021] At the same time, consider the fatigue reduction caused by workers' rest:
[0022]
[0023] Among them, μ is the fatigue recovery rate;
[0024] Consider the influence of fatigue degree on processing time:
[0025] pt′ i,j,k =pt i,j,k (1 + λδln(1 + f w )) (4)
[0026] Among them, pt i,j,k is the processing time of the jth process of workpiece i on machine k, and pt i ′ ,j,k is the processing time of the jth process of workpiece i on machine k considering the worker fatigue effect. δ is the influence degree of this process on the worker's processing time;
[0027] Consider the influence of workers with different skill levels on the processing time:
[0028]
[0029] where, \(pt\) i,j,k is the processing time of the \(j\)-th operation of workpiece \(i\) on machine \(k\) considering the worker fatigue effect and skill level, and \(T\) w is the skill level of worker \(w\);
[0030] Step 23, for workers, the unobservable and random human factor such as operation error, lack of interest and physical discomfort should also be considered. The final processing time composite formula is as follows:
[0031] \(pt\) i,j,k,w =\(pt''\) i,j,k,w +\(\omega\) i,j,k,w (6)
[0032] where, \(\omega\) i,j,k,w is the processing time fluctuation of the \(j\)-th operation of workpiece \(i\) on machine \(k\) processed by worker \(w\), which follows a normal distribution, and \(pt\) i,j,k,w is the processing time of the \(j\)-th operation of workpiece \(i\) on machine \(k\) processed by worker \(w\).
[0033] Furthermore, the design decision points in step 3 include: judging whether there are workpieces to be processed; if so, then checking whether the workpieces can be assigned to suitable processing machines; if possible, then judging whether the operation requires worker collaboration; if not, it is a scheduling decision point; if required, then further judging whether there is an available set of workers; if there are available workers, it is a scheduling decision point; otherwise, it is a non-scheduling decision point.
[0034] Furthermore, in step 4, the state space is designed based on the MAPPO scheduling framework. The workpiece agent and the worker agent work together as Actors. The workpiece agent generates a scheduling strategy according to the global state, and processes operation assignment, equipment selection and scheduling decisions; the worker agent optimizes worker assignment and task arrangement through the worker efficiency fluctuation model, combining factors such as fatigue and skill.
[0035] Furthermore, the state space is divided into global state features and local state features. The global features include 16 production attributes: the mean and standard deviation of the completion rate of all workpiece processing times; the mean and standard deviation of the completion rate of all workpiece operations; the mean and standard deviation of the remaining processing time of all workpieces; the mean and standard deviation of the current operation processing time of all workpieces; the mean and standard deviation of the utilization rate of all machines; the mean and standard deviation of the average processing rate of all machines; the mean and standard deviation of the utilization rate of all workers; the mean and standard deviation of the fatigue degree of all workers;
[0036] The local part includes 10 production attributes: the completion rate of the current workpiece processing time; the completion rate of the current workpiece process; the remaining processing time of the current workpiece; the process processing time of the current workpiece; the current machine utilization rate; the current average machining rate of the machine; the mean and standard deviation of the utilization rate of available workers; the mean and standard deviation of the fatigue degree of available workers.
[0037] Furthermore, the action space of the workpiece agent in step 5 includes 3 machine scheduling rules and 4 workpiece process scheduling rules, a total of 12 composite scheduling rules; the action space of the worker agent includes 6 worker scheduling rules.
[0038] Furthermore, in the design of the reward function in step 6, based on the ARSI training mechanism, the ARSI training mechanism includes two parts: autonomous reward and shared information. The autonomous reward includes the following:
[0039] The single-step reward calculation formula of the workpiece agent is as follows:
[0040]
[0041] Among them, r J represents the single-step reward function without considering human factors. pt i,j is the processing time of the j-th process of workpiece i, which is fixed and not affected by human factors; empty represents the idle time of the machine after the current scheduling, s represents the current state of the scheduling, s′ represents the state after the scheduling, m represents the machine, and M represents the set of machines;
[0042] The single-step reward calculation formula of the worker agent is as follows:
[0043]
[0044] Among them, r H represents the single-step reward function considering human factors. PT i,j is the processing time of the j-th process of workpiece i, and the influence of humans on the processing time needs to be considered. The range of this processing time is determined by the specified worker attributes.
[0045] Furthermore, the shared information part means that there is an information transmission process for each Actor during the execution and training phases. The execution phase specifically includes:
[0046] 1) Configure neural network layers for the workpiece agent and the worker agent. The hidden layer structure is 256 - 64 - 64 - 32. The dimensions of the input layer and the output layer are set according to the dimensions of the state space and the action space. The ReLU activation function is used in the hidden layer, and the Sigmoid function is used for mapping in the output layer to facilitate the subsequent calculation of the action selection probability using the Softmax function;
[0047] 2) Receive the workshop order information, extract the production data from the order and the workshop environment, and construct the production state space;
[0048] 3) Determine whether the current is a decision point. If so, enter the decision-making stage; the workpiece agent makes a scheduling decision based on the global state, and determines whether the participation of the worker agent is required; if required, the workpiece agent extracts local features, and the worker agent makes a decision based on the global and local states, and finally stores the decision data in the experience library;
[0049] 4) Determine whether the order is completed. If completed, stop execution; if not completed, return to the decision-making stage and continue execution until the task is completed.
[0050] Furthermore, the training stage specifically includes:
[0051] 1) Configure neural network layers for the workpiece agent and the worker agent. The hidden layer structure is 256-64-64-32, and the output layer dimension is 1; the ReLU activation function is used for the input of the middle hidden layer, and the Relu function is also used for the output layer;
[0052] 2) Calculate the advantage functions of the workpiece Actor and the worker Actor respectively: The generalized advantage function of the workpiece agent is calculated based on the time difference residual between the reward and the predicted reward, and the reward trajectory of the worker agent fills the missing rewards through the immediate rewards of the workpiece agent;
[0053] 3) Update the Actor and Critic parameters.
[0054] Compared with the prior art, the beneficial effects of the present invention are:
[0055] By introducing multi-agent deep reinforcement learning, the present invention significantly improves the accuracy and efficiency of workshop scheduling. In particular, by separately modeling each worker and considering factors such as fatigue and skills, the production process is effectively optimized. The ARSI training mechanism introduced in the present invention improves the cooperation efficiency between the workpiece agent and the worker agent. Through the self-reward mechanism, the workpiece agent and the worker agent can adjust the rewards according to their respective decisions and states, and more precisely optimize the scheduling strategy; through the shared information mechanism, the agents share global information during the training stage, improve the collaborative working ability, and further enhance the stability, learning efficiency and the ability to adapt to complex production environments of the system.
[0056] This method can be widely applied to various types of workshops, comprehensively consider various attributes of workers, so as to meet the complex and variable production scheduling requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1Shows the overall framework of the flexible job shop scheduling method based on the HCPS concept;
[0058] Figure 2 Shows the components of the worker efficiency fluctuation model;
[0059] Figure 3 Is the decision point judgment flowchart;
[0060] Figure 4 Is the flowchart of multi-agent collaboration in the workshop system;
[0061] Figure 5 Is the ARSI training mechanism flowchart.
[0062] Figure 6 Is the comparison result of different algorithms under static scheduling;
[0063] Figure 7 Is the comparison result of different algorithms under dynamic scheduling. Detailed implementation method
[0064] The detailed description of the present invention will be combined with the drawings and examples to show its technical solution. In the following embodiments, the implementation methods and operation steps of the present invention will be specifically described. It should be noted that the protection scope of the present invention is not limited to these examples.
[0065] The scheduling method of the present invention takes the flexible job shop as the research object, comprehensively considers multiple factors, and takes minimizing the production time as the optimization goal. By reasonably planning worker assignment, process processing and machine allocation, efficient production scheduling is achieved. As Figure 1 shown, this method includes the following steps:
[0066] Step 1, construct the target optimization model: taking minimizing the makespan as the goal, model the flexible job scheduling problem as a Markov decision process.
[0067] Among them, to minimize the makespan, the formula is as follows:
[0068]
[0069] Among them, C max represents the total completion time of all processing tasks, x represents the workpiece, n is the total number of workpieces, and ETx is the completion time of each workpiece.
[0070] Step 2, establish a worker efficiency fluctuation model for each worker separately, record the fatigue accumulation rate, fatigue recovery rate, skill level of each worker, as well as the working status and fatigue degree at each stage. Thus, effectively adjust the task assignment according to the actual situation of the workers, and further optimize the production scheduling and work efficiency, as Figure 2 shown:
[0071] First, initialize the fatigue accumulation rate, fatigue recovery rate, skill level of each worker, as well as the working state and fatigue degree at each stage.
[0072] Secondly, considering the fatigue accumulation of workers, the change in fatigue degree takes into account the fatigue accumulation caused by work. The fatigue accumulation calculation formula is as follows:
[0073]
[0074] Among them, f w (t i ) is the fatigue degree of worker w at the current moment; λ is the fatigue accumulation rate; t i represents the current moment; t i-1 represents the previous moment; i represents the index of time.
[0075] At the same time, the fatigue reduction caused by workers' rest is also considered:
[0076]
[0077] Among them, μ is the fatigue recovery rate.
[0078] Consider the influence of fatigue degree on processing time:
[0079] pt′ i,j,k =pt i,j,k (1 + λδln(1 + f w )) (4)
[0080] Among them, pt i,j,k is the processing time of the j-th process of workpiece i on machine k, pt′ i,j,k is the processing time of the j-th process of workpiece i on machine k considering the worker fatigue effect, and δ is the influence degree of this process on the worker's processing time.
[0081] Consider the influence of workers with different skill levels on processing time:
[0082]
[0083] Among them, pt″ i,j,k is the processing time of the j-th process of workpiece i on machine k considering the worker fatigue effect and skill level, and T w is the skill level of worker w.
[0084] Finally, for workers, unobservable and random human factor such as operation errors, lack of interest, and physical discomfort should also be considered. The final composite formula for processing time is as follows:
[0085] pt i,j,k,w= pt″ i,j,k,w + ò i,j,k,w (6)
[0086] where ò i,j,k,w is the processing time fluctuation of the j - th process of workpiece i processed by worker w on machine k, which follows a normal distribution, and pt i,j,k,w is the processing time of the j - th process of workpiece i processed by worker w on machine k.
[0087] Step 3: For the workshop, design decision points so as to carry out single - step scheduling operations, as Figure 3 shown.
[0088] The system determines whether there are workpieces to be processed. If there are, then it checks whether the workpieces can be assigned to suitable processing machines. If possible, it further determines whether the process requires worker cooperation. If not, it is a scheduling decision point; if so, it further determines whether there is an available set of workers. If there are available workers, it is a scheduling decision point; otherwise, it is not a scheduling decision point.
[0089] Step 4: Design the state space, Figure 4 which is the process for multi - agent cooperation based on MAPPO:
[0090] MAPPO is a multi - agent policy optimization algorithm extended from single - agent PPO and belongs to the Actor - Critic (AC) framework in reinforcement learning. The AC framework includes two core components: Actor and Critic, which perform different tasks respectively. The agent is the Actor, responsible for generating decision policies. The workpiece agent and the worker agent correspond to the workpiece Actor and the worker Actor respectively; the Critic evaluates the value of the current state and uses the evaluation result to help the Actor optimize the policy and improve the decision - making quality. The number and composition of the Critics are related to the training mechanism.
[0091] According to the complexity of the multi - agent environment, the production features are divided into global state features and local state features. Among them, the global features include 16 production attributes: the mean and standard deviation of the completion rate of all workpieces' processing time; the mean and standard deviation of the completion rate of all workpieces' processes; the mean and standard deviation of the remaining processing time of all workpieces; the mean and standard deviation of the current - process processing time of all workpieces; the mean and standard deviation of the utilization rate of all machines; the mean and standard deviation of the average processing rate of all machines; the mean and standard deviation of the utilization rate of all workers; the mean and standard deviation of the fatigue degree of all workers.
[0092] The local part includes 10 production attributes: the completion rate of the current workpiece processing time; the completion rate of the current workpiece process; the remaining processing time of the current workpiece; the process processing time of the current workpiece; the current machine utilization rate; the current average machine processing rate; the mean and standard deviation of the utilization rate of the available workers; the mean and standard deviation of the fatigue degree of the available workers.
[0093] Step Five, based on Figure 4 the MAPPO multi-agent collaboration process shown in
[0094] Design the action space for the workpiece agent:
[0095] The action space of the workpiece agent consists of 3 machine scheduling rules and 4 workpiece process scheduling rules, a total of 12 composite scheduling rules. The machine scheduling rules include: select the machine with the shortest current processing time; select the machine with the lowest utilization rate; select the machine with the longest remaining average processing time. The workpiece process scheduling rules include: select the workpiece with the shortest processing time; select the workpiece with the longest processing time; select the workpiece with the shortest remaining time outside the current operation; select the workpiece with the longest remaining time outside the current operation.
[0096] Step Six, based on the ARSI training mechanism, design the reward function:
[0097] ARSI is an autonomous reward and shared information mechanism, and its autonomous reward part means that each Actor has an independent reward mechanism. The autonomous reward includes the following: The single-step reward calculation formula for the workpiece Actor is as follows:
[0098]
[0099] where r J represents the single-step reward function without considering human factors, pt i,j is the processing time of the jth process of workpiece i, which is fixed and not affected by human factors; empty represents the idle time of the machine after the current scheduling, s represents the current state of the scheduling, s′ represents the state after the scheduling, m represents the machine, and M represents the set of machines.
[0100] The single-step reward calculation formula for the worker Actor is as follows:
[0101]
[0102] where r H represents the single-step reward function considering human factors, PT i,j$t_{ij}$ is the processing time of the $j$-th process of workpiece $i$. The influence of human factors on the processing time needs to be considered, and the range of this processing time is determined by the specified attributes of the processing worker.
[0103] Step 7: Use the ARSI training mechanism to train the model:
[0104] ARSI is an autonomous reward and shared information mechanism. The shared information part means that each Actor has an information transfer process during both the execution and training phases. During the execution phase, the workpiece Actor makes decisions based on the global state and determines whether the participation of the worker Actor is required. If so, local features are extracted according to the strategy made by the workpiece Actor, and then the worker Actor makes decisions based on the global and local states.
[0105] During the training phase, in order to improve the cooperation ability between Actors, the workpiece Critic and the worker Critic jointly share the global information. This information sharing helps to optimize the strategy, enabling Actors to work together better and improving the overall performance of the system.
[0106] As Figure 4 shown, the execution phase specifically includes:
[0107] First step: First, configure the neural network layers for the workpiece Actor and the worker Actor. The hidden layer structures of both Actors are 256-64-64-32. The dimensions of the input layer and the output layer depend on the dimensions of the state space and the action space. The ReLU activation function is used for the intermediate hidden layers to fit non-linear situations; the Sigmoid function is used for the output layer to map the output to the interval (-1,1), facilitating the subsequent calculation of the probability of action selection through the Softmax function.
[0108] Second step: When receiving a workshop order, extract production data from the order and the workshop environment, and then process this data to form a production state space.
[0109] Third step: Determine whether the current is a decision point. If so, enter the decision-making phase. The workpiece Actor makes decisions based on the global state and determines whether the participation of the worker Actor is required. If so, the workpiece Actor extracts local features, and the worker Actor makes decisions based on the global and local states, and finally stores the data in the experience library.
[0110] Fourth step: Determine whether the order is completed. If completed, stop the execution phase; otherwise, return to the third step and continue the execution.
[0111] During the training phase, to improve the collaboration ability among Actors, the workpiece Critic and the worker Critic jointly share global information. This information sharing helps optimize the policy, enabling the Actors to work together better and enhancing the overall system performance.
[0112] The specific training steps are as Figure 5 shown:
[0113] First step, first configure the neural network layers for the workpiece Critic and the worker Critic. The hidden layer structure of both Critics is 256-64-64-32. The dimension of the input layer depends on the state space dimension, and the dimension of the output layer is 1 for both. The ReLU activation function is used for the intermediate hidden layers to fit non-linear situations; the ReLU function is also used for the output layer, acting as a value function to evaluate the state of the Actor.
[0114] Second step, calculate the advantage functions of the workpiece Actor and the worker Actor:
[0115] For the workpiece Actor, calculate its generalized advantage function:
[0116]
[0117] where is the generalized advantage function, representing the advantage estimate for the workpiece Actor at time step t; is the time difference residual, representing the difference between the reward at the current time step t and the predicted reward at the next time step; γ is the discount factor, used to control the influence degree of future rewards on the current decision-making, and here it is taken as 0.95; λ is the generalized advantage parameter, used to balance the advantages and disadvantages of the temporal difference method and the Monte Carlo method, and here it is taken as 0.95, V J (S t ) is the state value, representing the expected return of the Actor in state S t at time step t, which is calculated by the workpiece Critic.
[0118] Since the worker Actor does not participate in decision-making at every scheduling point and its reward trajectory is usually discontinuous, the traditional generalized advantage function estimation method is not applicable. To solve this problem, the missing rewards can be filled by using the immediate rewards of the workpiece Actor to make the reward trajectory continuous. The calculation method of its total return reward function is as follows:
[0119]
[0120] where R H represents the total return reward of the worker Actor; Denote the single-step reward of the worker Actor at time step i, where i = 1, 2, …; Denote the missing single-step reward of the worker Actor at time step t; Denote the single-step reward filled by the workpiece Actor for the missing single-step reward of the worker Actor at time step t, where t represents the time step of the missing single-step reward of the worker actor; Denote the single-step reward obtained by the worker Actor at the last time step; PT k Denote the total processing duration on machine k; empty k Denote the total idle duration of machine k during the entire processing; C H Denote the completion time of the processing stage considering human factors; m is the number of machines. The final cumulative reward function is proportional to the maximum processing time, verifying the effectiveness of this reward function.
[0121] To estimate the state-action value function of the worker Actor, this study adopts the Monte Carlo method (MC). Directly using the immediate reward to calculate the state-action value function may lead to the problem of high variance. To solve this problem, the λ-return algorithm and the temporal difference (TD) method are introduced to balance the variance and bias. In this process, the state value V H (s t ) calculated by the worker Critic is not directly used, but the state value V J (s t ) calculated by the workpiece Critic is used to replace the immediate reward part missing for the worker Actor. The calculation formula is as follows:
[0122]
[0123] Among them, Denote the advantage estimation for the worker Actor at time step t; Q H Is the state-action value function of the worker Actor; γ is the discount factor, used to weigh the importance of future return rewards; V H (s t ) Represents the state value, indicating the expected return of the worker Actor in state S t at time step t, calculated by the worker Critic.
[0124] The third step is parameter update.
[0125] MAPPO adopts the idea of proximal policy optimization (PPO), and limits the policy update amplitude by clipping the objective function, thereby improving the stability of training. Specifically, the goal of the Actor policy update is to maximize the following objective function:
[0126]
[0127] Among them, r t (θ) is the probability ratio between the current policy and the old policy, is the estimated value of the advantage function, and ∈ is the clipping factor, which is used to limit the amplitude of the Actor policy update.
[0128] The final estimated calculation of the advantage function is as follows:
[0129]
[0130] Among them, is the estimated value of the advantage function of the workpiece Actor; is the estimated value of the advantage function of the worker Actor; is the final estimated value of the advantage function of the workpiece Actor; is the final estimated value of the advantage function of the worker Actor; c1 and c2 are weight factors, which are 0.9 and 0.1 respectively here.
[0131] Through the above formula, the policy parameters are updated:
[0132] Among them, θ represents the neural network parameters of the Actor; α represents the learning rate of the Actor; L clip represents the loss function of the Actor, represents the gradient of the derivative of θ.
[0133] The update of the Critic value function aims to minimize the mean square error between the predicted value and the actual return. Specifically, the loss function of the Critic value function is as follows:
[0134]
[0135] Among them, V φ (s t ) is the predicted value of the Critic value function under the current state S t , is the estimated value of the actual return.
[0136] Through the above formula, the parameters of the Critic value function are updated:
[0137] Among them, φ represents the neural network parameters of the Actor; β represents the learning rate of the Actor; L VF represents the loss function of the Actor.
[0138] To verify the effectiveness of the proposed algorithm, the present invention designs a comparative experiment that conforms to the scheduling rules. The rules are composed of the basic scheduling groups mentioned above: select three machine scheduling rules (MR1 - MR3: select the machine with the shortest current processing time; select the machine with the lowest utilization rate; select the machine with the longest remaining average processing time), two operation scheduling rules (JR1 - JR2: select the workpiece with the shortest processing time; select the workpiece with the shortest remaining time except for the current operation), and two worker scheduling rules (HR1 - HR2: select the worker with the shortest processing time; select the worker with the lowest fatigue degree). Eventually, 12 combinations of scheduling rules are formed: (1) MR1 + JR1 + HR1;
[0139] (2) MR1 + JR1 + HR2; (3) MR1 + JR2 + HR1; (4) MR1 + JR2 + HR2; (5) MR2 + JR1 + HR1; (6) MR2 + JR1 + HR2; (7) MR2 + JR2 + HR1; (8) MR2 + JR2 + HR2; (9) MR3 + JR1 + HR1; (10) MR3 + JR1 + HR2; (11) MR3 + JR2 + HR1; (12) MR3 + JR2 + HR2.
[0140] In the experiment, the cases are selected from real cases in the production workshop, and static scheduling experiments and dynamic scheduling experiments are respectively carried out. In the static scheduling experiment, the proposed algorithm and the composite scheduling rule respectively conduct scheduling planning for 8 groups of cases with different scales, repeat the experiment 20 times, record the maximum completion time each time, and take the average value as the final result; in the dynamic scheduling experiment, to simulate the scenario of receiving continuous orders in a real workshop, 10 groups of small-scale cases are sequentially sent to the workshop, and the proposed algorithm and the composite scheduling rule respectively conduct 40 repeated experiments, recording the maximum completion time each time.
[0141] As Figure 6 shown, in the static scheduling scenario, the proposed algorithm performs better than the composite scheduling rule. Specifically, it obtains the shortest processing time in all 8 cases, and in cases 3, 4, 5, and 6, it improves by 11.73%, 12.10, 11.51%, and 11.08% respectively compared to the second-best solution, fully demonstrating the good performance of the proposed algorithm. As Figure 7 shown, in the dynamic scheduling scenario, the advantage of the proposed algorithm is very prominent. It performs very well compared to the composite scheduling rule in reducing the completion time and time fluctuation. Especially in terms of reducing the completion time, it is reduced by 7.16% overall compared to the second-best solution.
[0142] The parts not described in the present invention are the same as or implemented by the prior art.
Claims
1. A scheduling method for an HCPS workshop system based on multi-agent deep reinforcement learning, characterized in that It includes the following steps: Step 1: Build an objective optimization model with the goal of minimizing the completion time. Step 2: Build a worker efficiency fluctuation model for each worker separately. Step 3: Model the HCPS workshop system as a Markov decision process and design decision points. Step 4: Design the state space. Step 5: Design the action space: Step 6: Design a reward function based on the ARSI training mechanism.
2. The scheduling method for an HCPS workshop system based on multi-agent deep reinforcement learning according to claim 1, characterized in that, The formula of the objective optimization model in Step 1 is as follows: Among them, C max represents the total completion time of all processing tasks, x represents the workpiece, n is the total number of workpieces, and ETx is the completion time of each workpiece.
3. A scheduling method for an HCPS workshop system based on multi-agent deep reinforcement learning according to claim 1, characterized in that The establishment of the worker efficiency fluctuation model in Step 2 specifically includes: Step 21: Initialize the worker fatigue accumulation rate, fatigue recovery rate, skill level of each worker, as well as the working status and fatigue degree at each stage. Step 22: Consider the worker's fatigue accumulation. The change in fatigue degree considers the fatigue accumulation caused by work. The fatigue accumulation calculation formula is as follows: Among them, f w (t i ) is the fatigue level of worker w at the current moment; λ is the fatigue accumulation rate; t i represents the current moment; t i-1 represents the previous moment; i represents the index of time; At the same time, consider the fatigue reduction caused by the worker's rest: Among them, μ is the fatigue recovery rate; Consider the influence of fatigue degree on the processing time: pt′ i,j,k = pt i,j,k (1 + λδln(1 + f w )) (4) Among them, \(pt_{ijk}\) i,j,k is the processing time of the \(j\)-th operation of workpiece \(i\) on machine \(k\). \(pt_{ijk}'\) i ′ ,j,k is the processing time of the \(j\)-th operation of workpiece \(i\) on machine \(k\) considering the worker fatigue effect. \(\delta\) is the influence degree of this operation on the worker's processing time; Consider the influence of workers with different skill levels on the processing time: where, pt″ i,j,k is the processing time of the j-th operation of workpiece i on machine k considering the worker fatigue effect and skill level, T w is the skill level of worker w; Step 23: For workers, unobservable and random human factors such as operation errors, lack of interest, and physical discomfort should also be considered. The final processing time composite formula is as follows: Among them, is the processing time fluctuation of the j-th process of workpiece i processed by worker w on machine k, which follows a normal distribution, pt i,j,k,w is the processing time of the j-th process of workpiece i processed by worker w on machine k.
4. The scheduling method for an HCPS workshop system based on multi-agent deep reinforcement learning according to claim 1, characterized in that The design of decision points in Step 3 includes: Determine whether there are workpieces to be processed; if so, then check whether the workpiece can be assigned to a suitable processing machine; if it can, then judge whether the process requires worker cooperation; if not, it is a scheduling decision point; if it does, then further judge whether there is an available set of workers; if there are available workers, it is a scheduling decision point; otherwise, it is a non-scheduling decision point.
5. A scheduling method for an HCPS workshop system based on multi-agent deep reinforcement learning according to claim 1, characterized in that In Step 4, the state space is designed based on the MAPPO scheduling framework. The workpiece agent and the worker agent work together as Actors. The workpiece agent generates a scheduling strategy based on the global state, dealing with process assignment, equipment selection, and scheduling decisions; the worker agent optimizes worker allocation and task arrangement through the worker efficiency fluctuation model, combining factors such as fatigue degree and skill.
6. The scheduling method for an HCPS workshop system based on multi-agent deep reinforcement learning according to claim 5, characterized in that, The state space is divided into global state features and local state features. The global features include 16 production attributes: the mean and standard deviation of the completion rate of all workpiece processing times; the mean and standard deviation of the completion rate of all workpiece processes; the mean and standard deviation of the remaining processing time of all workpieces; the mean and standard deviation of the current process processing time of all workpieces; the mean and standard deviation of the utilization rate of all machines; the mean and standard deviation of the average processing rate of all machines; the mean and standard deviation of the utilization rate of all workers; the mean and standard deviation of the fatigue degree of all workers; The local includes 10 production attributes: the completion rate of the current workpiece processing time; the completion rate of the current workpiece process; the remaining processing time of the current workpiece; the process processing time of the current workpiece; the utilization rate of the current machine; the average processing rate of the current machine; the mean and standard deviation of the utilization rate of available workers; the mean and standard deviation of the fatigue degree of available workers.
7. A scheduling method for an HCPS workshop system based on multi-agent deep reinforcement learning according to claim 1, characterized in that In Step 5, the action space of the workpiece agent includes 3 machine scheduling rules and 4 workpiece process scheduling rules, a total of 12 composite scheduling rules; the action space of the worker agent includes 6 worker scheduling rules.
8. The scheduling method for an HCPS workshop system based on multi-agent deep reinforcement learning according to claim 1, wherein In step 6, the design of the reward function is based on the ARSI training mechanism. The ARSI training mechanism consists of two parts: autonomous reward and shared information. The autonomous reward includes the following: The formula for calculating the single-step reward of the workpiece agent is as follows: Among them, r J represents the single-step reward function without considering human factors, and \(p_{t}\) i,j is the processing time of the \(j\)-th process of workpiece \(i\), which is fixed and not affected by human factors; empty represents the idle time of the machine after the current scheduling, \(s\) represents the current state of the scheduling, \(s'\) represents the state after the scheduling, \(m\) represents the machine, and \(M\) represents the set of machines; The formula for calculating the single-step reward of the worker agent is as follows: Among them, r H represents the single-step reward function considering human factors, and PT i,j is the processing time of the j-th process of workpiece i. The influence of humans on the processing time needs to be considered, and the range of this processing time is determined by the specified processing worker attributes.
9. A scheduling method for an HCPS workshop system based on multi-agent deep reinforcement learning according to claim 8, characterized in that, The shared information part indicates that there is a process of information transmission for each Actor during both the execution and training phases. The execution phase specifically includes: 1) Configure neural network layers for the workpiece agent and the worker agent. The hidden layer structure is 256-64-64-32 for both. The dimensions of the input layer and output layer are set according to the dimensions of the state space and action space. The ReLU activation function is used for the hidden layer, and the Sigmoid function is used for the output layer for subsequent calculation of the action selection probability using the Softmax function; 2) Receive workshop order information, extract production data from the order and workshop environment, and form a production state space; 3) Determine whether the current is a decision point. If so, enter the decision-making phase. The workpiece agent makes a scheduling decision based on the global state and determines whether the worker agent's participation is required. If required, the workpiece agent extracts local features, and the worker agent makes a decision based on the global and local states, and finally stores the decision data in the experience library; 4) Determine whether the order is completed. If completed, stop execution. If not completed, return to the decision-making phase and continue execution until the task is completed.
10. A scheduling method for an HCPS workshop system based on multi-agent deep reinforcement learning according to claim 9, characterized in that, The training phase specifically includes: 1) Configure neural network layers for the workpiece agent and the worker agent. The hidden layer structure is 256-64-64-32 for both, and the output layer dimension is 1. The ReLU activation function is used for the intermediate hidden layer output, and the Relu function is also used for the output layer; 2) Calculate the advantage functions of the workpiece Actor and the worker Actor respectively: The generalized advantage function of the workpiece agent is calculated based on the time difference residual between the reward and the predicted reward. The reward trajectory of the worker agent fills in the missing rewards through the immediate rewards of the workpiece agent; 3) Update the Actor and Critic parameters.
Citation Information
Cited By
Multi-agent collaborative dynamic target interception decision-making method based on reinforcement learning
CN121069790A
Human-machine cooperation assembly unit task intelligent scheduling method considering fatigue recovery
CN121436592A
Man-machine collaborative assembly line scheduling method and system based on reinforcement learning
CN121504047A
Multi-target reinforcement learning man-machine cooperation assembly task allocation method and system based on neighborhood parameter migration
CN121660326A
Man-machine collaborative dynamic scheduling system and method based on action mask and reward shaping MAPPO
CN122172756A