A green dynamic multi-objective scheduling method for flexible assembly workshops under the personnel learning effect

Through the dual-agent deep reinforcement learning method, the multi-objective scheduling conflicts in the flexible assembly workshop were resolved, the maximum completion time and energy consumption were optimized, and dynamic scheduling was achieved under the personnel learning effect and environmental constraints.

CN119739131BActive Publication Date: 2025-10-03CHENGDU UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510240011.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-10-03
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

Existing technologies find it difficult to effectively coordinate multi-objective scheduling conflicts in flexible assembly workshops, especially dynamic low-carbon scheduling problems under personnel learning effects and environmental factors, and existing algorithms cannot meet actual production needs.

Method used

A dual-agent structure and a two-layer deep reinforcement learning method are adopted, combined with immediate reward and delayed reward functions, to construct state space and action space, coordinate multi-objective scheduling conflicts, and optimize process-machine scheduling strategies.

Benefits of technology

Dynamic multi-objective scheduling of flexible assembly workshops under personnel learning effects and environmental constraints is achieved, the maximum completion time and total energy consumption of the processing process are optimized, and the real-time and effectiveness of scheduling are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119739131B_ABST
    Figure CN119739131B_ABST
Patent Text Reader

Abstract

The present invention discloses a green dynamic multi-objective scheduling method for a flexible assembly workshop under the human learning effect, and relates to the technical field of workshop scheduling. Based on the green dynamic multi-objective scheduling requirements of a flexible assembly workshop under the human learning effect, the present invention establishes a multi-objective dual-agent planning model. Under the dual resource constraints of machines and personnel, the model's dynamic multi-objective scheduling problem is solved based on a two-layer deep reinforcement learning framework method. The state space and action space are designed based on the workshop simulation environment, and a reward function combining immediate rewards and round rewards is combined. The agent interacts with the scheduling environment to obtain a more optimal scheduling rule at the scheduling point. Through the above method, the present invention improves the decision-making efficiency of manufacturing enterprises, can adaptively and quickly generate a more optimal solution, effectively reduce losses caused by delay time, and reduce energy consumption, while also having certain generalization and stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of workshop scheduling, and in particular to a green dynamic multi-objective scheduling method for a flexible assembly workshop under the personnel learning effect based on deep reinforcement learning. Background Art

[0002] In recent years, with the development of globalization and computer technology, many manufacturing companies have gradually shifted from the traditional job shop model to the assembly shop model, which can reduce material costs and improve factory production efficiency. Compared with the classic flexible job shop scheduling problem, the products scheduled in the low-carbon flexible assembly shop are composed of multiple modular parts. After each part is processed in the flexible job shop, it is transported to the assembly shop for assembly to complete the finished product. Therefore, the assembly job production model is more suitable for actual production environments and can better meet user customization needs. The field of deep reinforcement learning scheduling based on assembly shops is still in its early and mid-stages, making the design of scheduling solutions challenging. In addition, for multi-objective scheduling algorithms, many algorithm frameworks have the disadvantage of only optimizing a single objective at each rescheduling point, which cannot effectively coordinate conflicts between multiple objectives.

[0003] With the implementation and advancement of industrial manufacturing strategies, manufacturing systems are moving toward greater customization, and human-machine collaboration has become a mainstream development trend in intelligent manufacturing. In large-scale equipment manufacturing, complex welding and assembly processes make robotic automation difficult to achieve, making highly qualified workers a critical production resource. The learning effect of personnel is a significant factor influencing the scheduling problem in dual-resource flexible job shops. Therefore, shop managers must determine the appropriate allocation of efficiency between machines and workers. Furthermore, job shop scheduling is a key means for manufacturing companies to reduce carbon emissions. Researchers have made some contributions to green scheduling, but these contributions are primarily theoretical and fall short of practical scheduling requirements. For example, they fail to consider energy savings during actual machine operation, use green scheduling indicators to build models more aligned with actual production scheduling, or develop more effective algorithms for solving green scheduling problems.

[0004] In summary, flexible job shop scheduling is an NP-hard problem, and expanding the problem to include human and environmental factors further complicates it. The low-carbon, dynamic, multi-objective scheduling problem for flexible assembly job shops under dual resource constraints is a novel and significant topic. It represents a complex production environment encompassing both job shops and assembly plants, characterized by large scale, dual constraints, a complex and dynamic environment, and high dynamics. Furthermore, due to the dynamic nature of the scheduling process, unexpected events such as urgent insertions can occur during production. These incidents can impact the efficiency of previously generated scheduling plans, shifting the scheduling objective from optimality to rapid rationality. Summary of the Invention

[0005] This application provides a green dynamic multi-objective scheduling method for flexible assembly workshops under the personnel learning effect, which takes minimizing the maximum completion time and the total energy consumption of the processing process as the optimization goal, and solves the dynamic low-carbon multi-objective scheduling problem of flexible assembly workshops under dual resource constraints by using a dual-agent-based solution method;

[0006] According to the first aspect disclosed in the present application, the present application provides a green dynamic multi-objective scheduling method for a flexible assembly workshop under the personnel learning effect, comprising the following steps:

[0007] S1: Based on the green dynamic multi-objective scheduling requirements of flexible assembly workshops under the personnel learning effect, a green dynamic multi-objective scheduling planning model for flexible assembly workshops under the personnel learning effect is established;

[0008] S2: Using a dual-agent structure to solve a green dynamic multi-objective scheduling planning model for a flexible assembly workshop under the human learning effect, and building a two-layer deep reinforcement learning framework to coordinate and resolve multi-objective conflicts;

[0009] S3: Construct a state space and action space that matches the problem model and algorithm framework, and propose an immediate reward function and a delayed reward function.

[0010] In one feasible embodiment, step S1 establishes a green dynamic multi-objective scheduling planning model for a flexible assembly workshop under the human learning effect, based on the green dynamic multi-objective scheduling requirements of the flexible assembly workshop under the human learning effect, taking into account the maximum completion time and total energy consumption of multiple workpieces during the processing and assembly process. The model parameters are as follows:

[0011] (1) The objective function includes the function for calculating the total processing time and total energy consumption function :

[0012] The maximum completion time is the maximum end time of the last process of all workpieces. The function is composed of all workpieces. Maximum completion time calculate, is the total number of workpieces, Used to return the largest value from a set of values. Used to return the smallest value from a set of values;

[0013] ;

[0014] The total energy consumption of the processing process includes processing energy consumption and idle energy consumption;

[0015] Indicates processing energy consumption:

[0016] ;

[0017] in, Indicates workers Operating equipment Execution process The initial processing time, Represents workpiece No. process, Representation device Unit processing energy consumption, Indicates the judgment process Whether by workers On the device 0-1 decision variables executed on, Indicates workers Use machine Processing procedures The actual learning rate when Indicates the total number of devices, is the total number of personnel, For workpiece The total number of processes;

[0018] Indicates the idle energy consumption of computing equipment:

[0019] ;

[0020] in, Indicates the idle energy consumption of computing equipment, Representation device Unit idle energy consumption, Indicates the process By workers On the device The start time on Indicates the process By workers On the device The end time of Indicates the judgment process On the device Is the subsequent process on 0-1 decision variables, Indicates the judgment process Whether by workers On the device 0-1 decision variables executed on, For workpiece The total number of processes, For workpiece The total number of processes;

[0021] (2) A series of constraints include:

[0022] Limit a process to be operated by one person on one piece of equipment;

[0023] ;

[0024] The actual processing time of the process is equal to the end time minus the start time. By the workers On the device Upper operation process The initial processing time, the actual running processing time is obtained by the initial processing time and learning rate, By the workers On the device Upper operation process The end time of the operation, By the workers On the device Upper operation process The start time of the operation, Indicates workers Use machine Processing procedures The actual learning rate when

[0025] ;

[0026] Process Processing completion time Equal to the start processing time of the process plus the process processing time, equal to the worker Operating equipment Processing procedures End time: ;

[0027] The process of each workpiece must follow the priority order from front to back, that is, the process By workers In the machine Start time on Not less than the previous process of the same workpiece By workers In the machine End time of the run ;

[0028] ;

[0029] If you want to process different workpieces on one device, you must do it in sequence. 1, indicating a process On the device The subsequent process is , 1, indicating a process On the device The subsequent process is , there is only one of the two situations, process By workers In the machine The start time of the above operation is , end time is , process By workers In the machine The start time of the above operation is , end time is ;

[0030]

[0031] ;

[0032] Any machine can only operate one process at a time, that is, the worker In the machine The last step after processing Start time No less than workers On the same machine The previous process End time , is a positive number;

[0033] ;

[0034] Any person can only operate one process at a time. Indicates workers Operate the machine first Operate the machine again , Indicates workers Can operate the machine The start time of

[0035] ;

[0036] Worker Operating the machine Processing procedures Start time No less than the machine The end time of the previous process is equal to the worker In the machine The start time of the operation ;

[0037] ;

[0038] Worker Operate the machine again The interval before It is equivalent to the last time the machine was operated. End time To operate the machine again Start time The interval time between

[0039] ;

[0040] Worker Processing to The actual learning rate of the worker at the time of Number of times worked on related, Dynamically adjust the learning rate, is a worker The learning effect coefficient, It is the incompressible coefficient of the learning effect. The more times, the lower the learning rate, and vice versa.

[0041] ;

[0042] Assembly workpiece The first process The start time is no earlier than The end time of all predecessor jobs, Indicated by workers Operating the machine Processing procedures The start time, Represents workpiece The set of predecessor artifacts, Represents workpiece The last step, Indicated by workers operate Processing procedures End time;

[0043] .

[0044] In a feasible implementation, step S2 proposes a dual-agent framework to solve the green dynamic multi-objective scheduling planning model of a flexible assembly workshop under the human learning effect, and constructs a two-layer deep reinforcement learning framework to coordinate and resolve multi-objective conflicts;

[0045] A dual-agent approach is used to achieve multiple goals, designing a hierarchical multi-action space. Each space is managed by a DQN (DeepQ-Network)-based agent. When the framework is running, state information is first passed to the upper-layer DQN agent. After obtaining the output, the output and environmental state characteristics are then passed to the lower-layer DQN agent to obtain the final scheduling result for that time step.

[0046] The high-level DQN agent is a controller. The output result is used to determine the preference value of the optimization target of the current time step, guiding the selection of scheduling actions by the lower-level DQN agent, thereby affecting the scheduling of the entire environment state. Using the target preference value will better coordinate the conflicts between multiple targets, and will not make the lower-level action selection optimize only target 1 or target 2, but will be biased on the basis of balance. The low-level DQN agent receives the optimization preference value of the high-level DQN agent and uses it as the target parameter of the reward function. Then, through the input environment state information, it selects the current scheduling point according to the rules of the action space scheduling algorithm. The action is right;

[0047] The dual-agent real-time control workshop production process is as follows:

[0048] (1) High-level DQN

[0049] The high-level DQN agent acts as a controller. At each rescheduling point, it takes the state features as input and the Q value of each target as output. It then passes the output target probability value as a weight to the low-level DQN agent.

[0050] Target value in the DQN algorithm The calculation method is as follows:

[0051] ;

[0052] in, Is in state Next action After receiving the instant reward, is the discount factor, are the target network parameters, The target network For the next state Maximum Q value estimation;

[0053] The error calculation formula between the estimated value and the target value in the current state, that is, the loss function, is as follows:

[0054] ;

[0055] The agent is in state Next action The immediate reward obtained from the environment reflects the direct effect of the current action, is the discount factor, is the next state of the target network The maximum Q value estimate, Is the main network's current state and actions Q value estimation;

[0056] (2) Low-level PPO

[0057] The low-level PPO (Proximal Policy Optimization) agent acts as the executor, taking the state features and the output target of the high-level DQN agent as input, and the Q value of the workpiece-machine sequence pair as output. The workpiece-machine with the highest Q value is selected as the action selected for the current scheduling point;

[0058] The following formula is the objective function of the PPO algorithm, by finding the parameters θ Maximize this function:

[0059] ;

[0060] It is The objective function of PPO at the iteration is: yes t The state of the moment, yes t Actions taken at all times, In the current state Take action strategy, is the importance sampling ratio, which represents the probability ratio of the new strategy to the old strategy, and is used to measure the change of the new strategy compared to the old strategy. Indicates limiting the importance sampling ratio to between, is a parameter that limits the range; It is The strategy parameters at the iteration, is the advantage function, It is The advantage function at the iteration, To evaluate the status Take action the pros and cons of; It means obtaining the minimum value among a set of target values, which can effectively prevent excessive deviation from the old strategy when the strategy is updated;

[0061] The decision-making part of the algorithm uses the Actor-Critic algorithm, which is divided into a Critic network and an Actor Policy Network. The former inputs the environment state and action to evaluate the value of taking the action in the current state. The latter inputs the current state and outputs the probability distribution of the action. The Critic network then uses it to evaluate the quality of the action and adjust the policy to improve the expected return of future actions.

[0062] The Critic network uses the mean square error loss function, that is, ; is the total number of samples, Corresponding to samples, is the mean square error loss function value obtained by the Critic network, and Calculate the cumulative reward value for each sample Comparing the status with the Critic network and actions Q-value estimation The mean square error between them promotes the Critic network to quickly correct the error, and finally The sum of the errors of the samples is averaged to obtain ;

[0063] The optimization goals of the Actor network are as follows: ; is the policy function, Is the current state Take action The logarithmic probability of the policy function is used to measure the current policy in the state Select Action possibility; Is the advantage function advantage estimate, indicating that in this state Take action Compared with the average strategy, if it is greater than 0, it means it is better, then the probability of selecting the action is increased; if it is less than 0, it means it is worse, then the probability of selecting the action is reduced. The error of each sample is summed and averaged to get the loss function value of the Actor network .

[0064] In a feasible implementation, step S3 constructs a state space and action space that match the problem model and algorithm framework, and proposes an immediate reward function and a delayed reward function;

[0065] (1) State space

[0066] The production system for low-carbon flexible assembly workshop scheduling consists of three elements: workpiece, processing machine, and personnel. The global state characteristics include the three elements of workpiece characteristics, processing machine characteristics, and personnel characteristics.

[0067] Workpiece characteristic indicators include actual completion time, remaining processing time, expected completion time and average completion time, as well as the predecessor-successor relationship between workpieces to indicate assembly operations; processing machine characteristic indicators include machine utilization rate and machine completion time; personnel characteristic indicators include worker initial learning rate, learning rate, and forgetting curve;

[0068] Rescheduling point or decision point Average utilization of all devices on the , is the total number of devices:

[0069] ;

[0070] in, Indicates that at a rescheduling point or decision point On the device Equipment utilization rate;

[0071] ;

[0072] Representation device Completion time of the last process;

[0073] Rescheduling point or decision point The standard deviation of equipment utilization :

[0074] ;

[0075] All artifacts at rescheduling points or decision points Average completion rate :

[0076] ;

[0077] Indicates that at a rescheduling point or decision point Loading workpiece The number of completed processes, Indicates the Number of processes per workpiece;

[0078] All processes are at the rescheduling point or decision point Average completion rate , is the total number of artifacts:

[0079] ;

[0080] in, Indicates that at a rescheduling point or decision point Loading workpiece completion rate;

[0081] ;

[0082] At a rescheduling point or decision point Estimated maximum completion time for loading workpieces :

[0083] ;

[0084] Indicates that at a rescheduling point or decision point Loading workpiece Completed process set, Indicates the process The running time, Indicates the process On the device 0-1 decision variables for the above operations;

[0085] Workpiece Actual running time of completed operations Determined by the personnel learning rate:

[0086] ;

[0087] Personnel learning rate The number of times personnel learn And the forgetting curve yields:

[0088] ;

[0089] At a rescheduling point or decision point Energy consumption indicators for completed processes :

[0090] ;

[0091] in, Indicates that at a rescheduling point or decision point The actual energy consumption of the completed process, Indicates that at a rescheduling point or decision point The minimum energy consumption required to complete the process, Indicates that at a rescheduling point or decision point The median energy consumption required to complete the process;

[0092] ;

[0093] ;

[0094] ;

[0095] in, Indicates processing steps The operating energy consumption generated by Indicates the process Start time to the same device Previous process Idle energy consumption at the end time, Indicates the process used to execute Processing equipment set, Indicates that at a rescheduling point or decision point The maximum energy consumption required to complete the process, Indicates the process used to execute Processing personnel set, Indicates the process By workers On the device End time on;

[0096] ;

[0097] ;

[0098] (2) Action Space

[0099] The high-level DQN agent calculates the probability of the target focus based on the current state, and the output indicates the degree of preference for the two targets in the current state;

[0100] The low-level PPO agent generates the optimal process-machine action pair by combining the high-level target ratio and the current environment state. It constructs a heterogeneous graph neural network (HGNN), uses a multi-layer perceptron (MLP) to embed operation nodes and machine nodes, and builds a graph attention network (GAT) model to learn using the attention mechanism. It then filters out operation-machine node pairs that do not meet the scheduling conditions. Finally, it inputs the Actor network to obtain the action probability vector. When selecting action a, it uses ε-Greedy to add randomness to the exploration:

[0101] ;

[0102] is the strategy function of the Actor network, is a deterministic strategy, which means choosing The biggest move, It is a random strategy, calculated based on the Actor network Probability distribution randomly selects actions;

[0103] (3) Reward Function

[0104] The reward function includes immediate rewards and delayed rewards;

[0105] Instant Rewards : Instant rewards include economic indicators and energy consumption indicators;

[0106] Calculating economic indicators , based on the average equipment utilization , standard deviation of equipment utilization , the expected maximum completion time during training To calculate, Indicates the current time point, Indicates the next rescheduling point:

[0107]

[0108]

[0109]

[0110] Calculate energy consumption indicators , according to the energy consumption index , the minimum total energy consumption during training and current total energy consumption To calculate:

[0111]

[0112]

[0113] The weighted sum of economic indicators and energy consumption indicators is used as the immediate reward, and the parameters Used to balance economic indicators and energy consumption indicators, where the parameters and It is the target optimization probability ratio passed from the high-level DQN to the lower layer, but it is necessary to keep the positive and negative values ​​of the reward specified by the state without being affected by the weight parameters:

[0114] ;

[0115] Delayed Rewards :

[0116] The delay reward is a negative value; the total completion time is generated after the lower-level agent selects the action pair and completes the scheduling process in the workshop environment and total energy consumption ,parameter and are the sum of the weighted probabilities of the two objectives in a complete scheduling process;

[0117] ;

[0118] The larger the total completion time and total energy consumption, the greater the penalty the environment gives back to the agent. The experimental scheduling results are presented using a Gantt chart.

[0119] Compared with the prior art, this application has the following beneficial effects:

[0120] The present invention uses a dual-agent structure to solve the green dynamic multi-objective scheduling problem of a flexible assembly workshop under the effect of personnel learning, and uses two-layer deep reinforcement learning and a reward function that combines immediate rewards and delayed rewards to coordinately solve the problem of multi-objective conflicts; the method of the present invention consists of two-layer deep reinforcement learning, the high layer combines the learning ability of deep convolutional neural networks with the decision-making ability of reinforcement learning, and the low layer adopts a policy optimization algorithm based on the high-layer results to obtain a better scheduling strategy to solve the scheduling problem, using the real-time production environment as the state space, using machine and personnel dual resource constraints, and obtaining a better real-time scheduling strategy based on real-time production environment information, achieving a balance between solution quality and algorithm dynamics; in addition, a reward function that combines immediate rewards and cumulative rewards is designed in the framework, and the agent continuously interacts with the environment to obtain the optimal scheduling action for each rescheduling point or decision point. BRIEF DESCRIPTION OF THE DRAWINGS

[0121] Figure 1 This is a low-carbon flexible assembly workshop scheduling architecture diagram;

[0122] Figure 2 It is the flow chart of the algorithm for generating the optimal process-machine action pair;

[0123] Figure 3 It is the personnel learning effect curve;

[0124] Figure 4 This is a Gantt chart of the experimental scheduling results. DETAILED DESCRIPTION

[0125] A dual-agent structure is used to solve a green dynamic multi-objective scheduling problem in a flexible assembly workshop under the effect of personnel learning. A double-layer deep reinforcement learning is used to coordinate and solve multi-objective conflicts. A reward function combining immediate rewards and delayed rewards is adopted. The low-carbon flexible assembly workshop architecture is as follows: Figure 1 As shown, data information such as the environmental state is input into the high-level intelligent agent to obtain the probability weight, that is, the target preference value. The weight is then passed to the low-level intelligent agent, and the optimal process-machine action pair is obtained by combining the data information. After selecting the personnel, the action scheduling is completed. The above operation is repeated until the scheduling of all workpieces in the algorithm is completed.

[0126] 1. Mathematical model

[0127] Considering the maximum completion time and maximum energy consumption of multiple workpieces during machining and assembly operations, a multi-objective mathematical programming model was established. The model parameters are shown in the following table:

[0128]

[0129] (1) The objective function includes the function for calculating the total processing time and total energy consumption function :

[0130] The maximum completion time is the maximum end time of the last process of all workpieces. The function is composed of all workpieces. Maximum completion time calculate, is the total number of workpieces, It represents the maximum completion time of all workpieces, that is, the total processing time of workpiece scheduling. Used to return the largest value from a set of values. Used to return the smallest value from a set of values;

[0131] ;

[0132] The total energy consumption of the processing process includes processing energy consumption and idle energy consumption;

[0133] ;

[0134] Indicates processing energy consumption:

[0135] ;

[0136] in, Indicates workers Operating equipment Execution process The initial processing time, Represents workpiece No. process, Representation device Unit processing energy consumption, Indicates the judgment process Whether by workers On the device 0-1 decision variables executed on, Indicates workers Use machine Processing procedures The actual learning rate when Indicates the total number of devices, is the total number of personnel, For workpiece The total number of processes;

[0137] Indicates the idle energy consumption of computing equipment:

[0138] ;

[0139] in, Indicates the idle energy consumption of computing equipment, Representation device Unit idle energy consumption, Indicates the process By workers On the device The start time on Indicates the process By workers On the device The end time of Indicates the judgment process On the device Is the subsequent process on 0-1 decision variables, Indicates the judgment process Whether by workers On the device 0-1 decision variables executed on, For workpiece The total number of processes, For workpiece The total number of processes;

[0140] (2) A series of constraints include:

[0141] Limit a process to be operated by one person on one piece of equipment;

[0142] ;

[0143] The actual processing time of the process is equal to the end time minus the start time. By the workers On the device Upper operation process The initial processing time, the actual running processing time is obtained by the initial processing time and learning rate, By the workers On the device Upper operation process The end time of the operation, By the workers On the device Upper operation process The start time of the operation, Indicates workers Use machine Processing procedures The actual learning rate when

[0144] ;

[0145] Process Processing completion time Equal to the start processing time of the process plus the process processing time, equal to the worker Operating equipment Processing procedures End time:

[0146] ;

[0147] The process of each workpiece must follow the priority order from front to back, that is, the process By workers In the machine Start time on Not less than the previous process of the same workpiece By workers In the machine End time of the run ;

[0148] ;

[0149] If you want to process different workpieces on one device, you must do it in sequence. 1, indicating a process On the device The subsequent process is , 1, indicating a process On the device The subsequent process is , there is only one of the two situations, process By workers In the machine The start time of the above operation is , end time is , process By workers In the machine The start time of the above operation is , end time is ;

[0150]

[0151] ;

[0152] Any machine can only operate one process at a time, that is, the worker In the machine The last step after processing Start time No less than workers On the same machine The previous process End time , is a positive number;

[0153] ;

[0154] Any person can only operate one process at a time. Indicates workers Operate the machine first Operate the machine again , Indicates workers Can operate the machine The start time of

[0155] ;

[0156] Worker Operating the machine Processing procedures Start time No less than the machine The end time of the previous process is equal to the worker In the machine The start time of the operation :

[0157] ;

[0158] Worker Operate the machine again The interval before It is equivalent to the last time the machine was operated. End time To operate the machine again Start time The interval time between

[0159] ;

[0160] Worker Processing to The actual learning rate of the worker at the time of Number of times worked on related, Dynamically adjust the learning rate, is a worker The learning effect coefficient, It is the incompressible coefficient of the learning effect. The more times, the lower the learning rate, and vice versa.

[0161] ;

[0162] Assembly workpiece The first process The start time is no earlier than The end time of all predecessor jobs, Indicated by workers Operating the machine Processing procedures The start time, Represents workpiece The set of predecessor artifacts, Represents workpiece The last step, Indicated by workers operate Processing procedures End time;

[0163] .

[0164] 2. Dual Agents

[0165] A dual-agent approach is used to achieve multiple goals, designing a hierarchical multi-action space. Each space is managed by a DQN (DeepQ-Network)-based agent. When the framework is running, state information is first passed to the upper-layer DQN agent. After obtaining the output, the output and environmental state characteristics are then passed to the lower-layer DQN agent to obtain the final scheduling result for that time step.

[0166] The high-level DQN agent is a controller. The output result is used to determine the preference value of the optimization target of the current time step, guiding the selection of scheduling actions by the lower-level DQN agent, thereby affecting the scheduling of the entire environment state. Using the target preference value will better coordinate the conflicts between multiple targets, and will not make the lower-level action selection optimize only target 1 or target 2, but will be biased on the basis of balance. The low-level DQN agent receives the optimization preference value of the high-level DQN agent and uses it as the target parameter of the reward function. Then, through the input environment state information, it selects the current scheduling point according to the rules of the action space scheduling algorithm. The action is right;

[0167] The dual-agent real-time control workshop production process is as follows:

[0168] (1) High-level DQN

[0169] The high-level DQN agent acts as a controller. At each rescheduling point, it takes the state features as input and the Q value of each target as output. It then passes the output target probability value as a weight to the low-level DQN agent.

[0170] Target value in the DQN algorithm The calculation method is as follows:

[0171] ;

[0172] in, Is in state Next action After receiving the instant reward, is the discount factor, are the target network parameters, The target network For the next state Maximum Q value estimation;

[0173] The error calculation formula between the estimated value and the target value in the current state, that is, the loss function, is as follows:

[0174] ;

[0175] The agent is in state Next action The immediate reward obtained from the environment reflects the direct effect of the current action, is the discount factor, is the next state of the target network The maximum Q value estimate, Is the main network's current state and actions Q value estimation;

[0176] (2) Low-level PPO

[0177] The low-level PPO (Proximal Policy Optimization) agent acts as the executor, taking the state features and the output target of the high-level DQN agent as input, and the Q value of the workpiece-machine sequence pair as output. The workpiece-machine with the highest Q value is selected as the action selected for the current scheduling point;

[0178] The following formula is the objective function of the PPO algorithm, by finding the parameters θ Maximize this function:

[0179] ;

[0180] It is The objective function of PPO at the iteration is: yes t The state of the moment, yes t Actions taken at all times, In the current state Take action strategy, is the importance sampling ratio, which represents the probability ratio of the new strategy to the old strategy, and is used to measure the change of the new strategy compared to the old strategy. Indicates limiting the importance sampling ratio to between, is a parameter that limits the range; It is The strategy parameters at the iteration, is the advantage function, It is The advantage function at the iteration, To evaluate the status Take action The pros and cons of It means obtaining the minimum value among a set of target values, which can effectively prevent excessive deviation from the old strategy when the strategy is updated;

[0181] The decision-making part of the algorithm uses the Actor-Critic algorithm, which is divided into a Critic network and an Actor Policy Network. The former inputs the environment state and action to evaluate the value of taking the action in the current state. The latter inputs the current state and outputs the probability distribution of the action. The Critic network then uses it to evaluate the quality of the action and adjust the policy to improve the expected return of future actions.

[0182] The Critic network uses the mean square error loss function, that is, ; is the total number of samples, Corresponding to samples, is the mean square error loss function value obtained by the Critic network, and Calculate the cumulative reward value for each sample Comparing the status with the Critic network and actions Q-value estimation The mean square error between them promotes the Critic network to quickly correct the error, and finally The sum of the errors of the samples is averaged to obtain ;

[0183] The optimization goals of the Actor network are as follows: ; is the policy function, Is the current state Take action The logarithmic probability of the policy function is used to measure the current policy in the state Select Action possibility; Is the advantage function advantage estimate, indicating that in this state Take action Compared with the average strategy, if it is greater than 0, it means it is better, then the probability of selecting the action is increased; if it is less than 0, it means it is worse, then the probability of selecting the action is reduced. The error of each sample is summed and averaged to get the loss function value of the Actor network .

[0184] The low-level PPO algorithm process is as follows:

[0185] Preparation: Critic network learning rate and network parameters , Actor network learning rate and network parameters ;

[0186] 1) Standardize current status information;

[0187] 2) Heterogeneous neural networks In the iteration, the action pairs that do not meet the conditions are filtered out and input into the Actor network to obtain the action probability;

[0188] Select actions based on rules , calculate the reward value ;

[0189] Will Store in cache;

[0190] 3) Round strategy optimization:

[0191] From the interaction between the agent and the environment Select from collection Group;

[0192] Get action evaluation through the Critic network;

[0193] 4) The old strategy gets the new strategy:

[0194] Calculate the network loss value: ,

[0195] Update network parameters: , .

[0196] 3. State Space

[0197] The production system for low-carbon flexible assembly workshop scheduling consists of three elements: workpiece, processing machine, and personnel. The global state characteristics include the three elements of workpiece characteristics, processing machine characteristics, and personnel characteristics.

[0198] Workpiece characteristic indicators include actual completion time, remaining processing time, expected completion time and average completion time, as well as the predecessor-successor relationship between workpieces to indicate assembly operations; processing machine characteristic indicators include machine utilization rate and machine completion time; personnel characteristic indicators include worker initial learning rate, learning rate, and forgetting curve;

[0199] Rescheduling point or decision point Average utilization of all devices on the , is the total number of devices:

[0200] ;

[0201] in, Indicates that at a rescheduling point or decision point On the device Equipment utilization rate;

[0202] ;

[0203] Representation device Completion time of the last process;

[0204] At a rescheduling point or decision point The standard deviation of equipment utilization :

[0205] ;

[0206] All artifacts at rescheduling points or decision points Average completion rate :

[0207] ;

[0208] Indicates that at a rescheduling point or decision point Loading workpiece The number of completed processes, Indicates the Number of processes per workpiece;

[0209] All processes are at the rescheduling point or decision point Average completion rate , is the total number of artifacts:

[0210] ;

[0211] in, Indicates that at a rescheduling point or decision point Loading workpiece completion rate;

[0212] ;

[0213] At a rescheduling point or decision point Estimated maximum completion time for loading workpieces :

[0214] ;

[0215] Indicates that at a rescheduling point or decision point Loading workpiece Completed process set, Indicates the process The running time, Indicates the process On the device 0-1 decision variables for the above operations;

[0216] Workpiece Actual running time of completed operations Determined by the personnel learning rate, as follows Figure 3 As shown, the blue lines represent workers The learning curve of workers Operating equipment The vertical axis represents the actual running time of the process. Because the running time of each process on different equipment is inconsistent, the vertical axis scale cannot be expressed in actual values. Point 1 is the worker On the device Upper operation process Initial running time, workers during continuous learning On the device Upper operation process The actual running time will decrease along the learning curve, passing through points 2 and 3. At point 3, the worker Left the device That is, forgetting occurs and returning to the previous n The learning rate at scheduling point 2, continuing to learn will reduce the actual running time along the learning curve of the blue line. The red line represents the worker There is a learning lower limit, that is, the actual running time will not be 0:

[0217] ;

[0218] Personnel learning rate The number of times personnel learn And the forgetting curve:

[0219] ;

[0220] Rescheduling point or decision point Energy consumption indicators for completing the process :

[0221] ;

[0222] in, Indicates that at a rescheduling point or decision point The actual energy consumption of the completed process, Indicates that at a rescheduling point or decision point The minimum energy consumption required to complete the process, Indicates that at a rescheduling point or decision point The median energy consumption required to complete the process;

[0223] ;

[0224] ;

[0225] ;

[0226] in, Indicates processing steps The operating energy consumption generated by Indicates the process Start time to the same device Previous process Idle energy consumption at the end time, Indicates the process used to execute Processing equipment set, Indicates that at a rescheduling point or decision point The maximum energy consumption required to complete the process, Indicates the process used to execute Processing personnel set, Indicates the process By workers On the device End time on;

[0227] ;

[0228] ;

[0229] 4. Action Space

[0230] The high-level DQN agent outputs the target emphasis probability ratio based on the current state, and the output result indicates the preference degree for the two targets in the current state;

[0231] The low-level PPO agent generates the optimal process-machine action pair by combining the high-level target ratio and the current environment state. It constructs a heterogeneous graph neural network (HGNN), uses a multi-layer perceptron (MLP) to embed operation nodes and machine nodes, and builds a graph attention network (GAT) model to learn using the attention mechanism. It then filters out operation-machine node pairs that do not meet the scheduling conditions. Finally, it inputs the Actor network to obtain the action probability vector. When selecting action a, it uses ε-Greedy to add randomness to the exploration:

[0232] ;

[0233] is the strategy function of the Actor network, is a deterministic strategy, which means choosing The biggest move, It is a random strategy, calculated based on the Actor network Probability distribution randomly selects actions;

[0234] The low-level algorithm action selection process is as follows Figure 2As shown in the figure, after standardizing the current state information, the operation nodes and machine nodes are loaded into the heterogeneous graph neural network using a multi-layer perceptron and graph attention network model. After L rounds of iterations, the data is pooled and filtered and then passed to the actor network to obtain the action probabilities of all process-machine action pairs.

[0235] 5. Reward Function

[0236] The reward function includes immediate rewards and delayed rewards;

[0237] Instant Rewards : Instant rewards include economic indicators and energy consumption indicators;

[0238] Calculating economic indicators , based on the average equipment utilization , standard deviation of equipment utilization , the expected maximum completion time during training To calculate, Indicates the current time point, Indicates the next rescheduling point:

[0239]

[0240]

[0241]

[0242] Calculate energy consumption indicators , according to the energy consumption index , the minimum total energy consumption during training and current total energy consumption To calculate:

[0243]

[0244]

[0245] The weighted sum of economic indicators and energy consumption indicators is used as the immediate reward, and the parameters Used to balance economic indicators and energy consumption indicators, where the parameters and It is the target optimization probability ratio passed from the high-level DQN to the lower layer, but it is necessary to keep the positive and negative values ​​of the reward specified by the state without being affected by the weight parameters:

[0246] ;

[0247] Delayed Rewards :

[0248] The delay reward is a negative value; the total completion time is generated after the lower-level agent selects the action pair and completes the scheduling process in the workshop environment and total energy consumption ,parameter and are the sum of the weighted probabilities of the two objectives in a complete scheduling process;

[0249] ;

[0250] The larger the total completion time and total energy consumption, the greater the penalty the environment gives back to the agent. The experimental scheduling results are presented using a Gantt chart.

[0251] The Gantt chart results of the successful experimental assembly task scheduling based on the dual-agent deep reinforcement learning structure are as follows: Figure 4 As shown in the figure, the scheduling sequence obtained by the algorithm is used to complete the scheduling of the assembly workpiece under the dual resource constraints of machines and workers. Indicates different workpieces, each workpiece has a different color. Indicates different processes of the same workpiece, displayed in the same color. to The operations of five machines to complete all workpieces in the scheduled sequence are displayed on the timeline, achieving multi-objective optimization of minimizing the maximum completion time and the total energy consumption of the machining process.

[0252] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A green dynamic multi-objective scheduling method for a flexible assembly workshop under the personnel learning effect, characterized by: The following steps are involved: S1: Based on the green dynamic multi-objective scheduling requirements of flexible assembly workshops under the personnel learning effect, a green dynamic multi-objective scheduling planning model for flexible assembly workshops under the personnel learning effect is established; S2: Using a dual-agent structure to solve a green dynamic multi-objective scheduling planning model for a flexible assembly workshop under the human learning effect, and building a two-layer deep reinforcement learning framework to coordinate and resolve multi-objective conflicts; The high-level DQN agent acts as a controller. At each rescheduling point, it takes the state features as input and the Q value of each target as output. It then passes the output target probability value as a weight to the low-level DQN agent. The low-level PPO (Proximal Policy Optimization) agent acts as the executor, taking the state features and the output target of the high-level DQN agent as input, and the Q value of the workpiece-machine sequence pair as output. The workpiece-machine with the highest Q value is selected as the action selected for the current scheduling point; S3: Construct a state space and action space that matches the problem model and algorithm framework, and propose an immediate reward function and a delayed reward function; The weighted sum of economic indicators and energy consumption indicators is used as the immediate reward, and the parameters Used to balance economic indicators and energy consumption indicators, where the parameters and It is the target optimization probability ratio passed from the high-level DQN to the lower layer, but the positive and negative values ​​of the reward need to be specified by the state without being affected by the weight parameters.

2. The green dynamic multi-objective scheduling method for a flexible assembly workshop under the personnel learning effect according to claim 1 is characterized in that: In step S1, a green dynamic multi-objective scheduling planning model for a flexible assembly workshop under the human learning effect is established based on the green dynamic multi-objective scheduling requirements of the flexible assembly workshop under the human learning effect, taking into account the maximum completion time and total energy consumption of multiple workpieces during the processing and assembly process. The model parameters are as follows: (1) The objective function includes the function for calculating the total processing time and total energy consumption function : The maximum completion time is the maximum end time of the last process of all workpieces. The function is composed of all workpieces. Maximum completion time calculate, is the total number of workpieces, Used to return the largest value from a set of values. Used to return the smallest value from a set of values; ; The total energy consumption of the processing process includes processing energy consumption and idle energy consumption; ; Indicates processing energy consumption: ; in, Indicates workers Operating equipment Execution process The initial processing time, Represents workpiece No. process, Representation device Unit processing energy consumption, Indicates the judgment process Whether by workers On the device 0-1 decision variables executed on, Indicates workers Use machine Processing procedures The actual learning rate when Indicates the total number of devices, is the total number of personnel, For workpiece The total number of processes; Indicates the idle energy consumption of computing equipment: ; in, Indicates the idle energy consumption of computing equipment, Representation device Unit idle energy consumption, Indicates the process By workers On the device The start time on Indicates the process By workers On the device The end time of Indicates the judgment process On the device Is the subsequent process on 0-1 decision variables, Indicates the judgment process Whether by workers On the device 0-1 decision variables executed on, For workpiece The total number of processes, For workpiece The total number of processes; (2) A series of constraints include: Limit a process to be operated by one person on one piece of equipment; ; The actual processing time of the process is equal to the end time minus the start time. By the workers On the device Upper operation process The initial processing time, the actual running processing time is obtained by the initial processing time and learning rate, By the workers On the device Upper operation process The end time of the operation, By the workers On the device Upper operation process The start time of the operation, Indicates workers Use machine Processing procedures The actual learning rate when ; Process Processing completion time Equal to the start processing time of the process plus the process processing time, equal to the worker Operating equipment Processing procedures End time: ; The process of each workpiece must follow the priority order from front to back, that is, the process By workers In the machine Start time on Not less than the previous process of the same workpiece By workers In the machine End time of the run ; ; If you want to process different workpieces on one device, you must do it in sequence. 1, indicating a process On the device The subsequent process is , 1, indicating a process On the device The subsequent process is , there is only one of the two situations, process By workers In the machine The start time of the above operation is , end time is , process By workers In the machine The start time of the above operation is , end time is ; ; Any machine can only operate one process at a time, that is, the worker In the machine The last step after processing Start time No less than workers On the same machine The previous process End time , is a positive number; ; Any person can only operate one process at a time. Indicates workers Operate the machine first Operate the machine again , Indicates workers Can operate the machine The start time of ; Worker Operating the machine Processing procedures Start time No less than the machine The end time of the previous process is equal to the worker In the machine The start time of the operation ; ; Worker Operate the machine again The interval before It is equivalent to the last time the machine was operated. End time To operate the machine again Start time The interval time between ; Worker Processing to The actual learning rate of the worker at the time of Number of times worked on related, Dynamically adjust the learning rate, is a worker The learning effect coefficient, It is the incompressible coefficient of the learning effect. The more times, the lower the learning rate, and vice versa. ; Assembly workpiece The first process The start time is no earlier than The end time of all predecessor jobs, Indicated by workers Operating the machine Processing procedures The start time, Represents workpiece The set of predecessor artifacts, Represents workpiece The last step, Indicated by workers operate Processing procedures End time; 。 3. The green dynamic multi-objective scheduling method for a flexible assembly workshop under the personnel learning effect according to claim 1 is characterized in that: In step S2, a dual-agent framework is proposed to solve the green dynamic multi-objective scheduling planning model of the flexible assembly workshop under the human learning effect, and a two-layer deep reinforcement learning framework is constructed to coordinate and resolve multi-objective conflicts; A dual-agent approach is used to achieve multiple goals, designing a hierarchical multi-action space. Each space is managed by a DQN (Deep Q-Network)-based agent. When the framework is running, state information is first passed to the upper-layer DQN agent. After obtaining the output, the output and environmental state characteristics are then passed to the lower-layer PPO agent to obtain the final scheduling result for that time step. The high-level DQN agent is a controller. The output result is used to determine the preference value of the optimization target at the current time step, guiding the selection of scheduling actions by the lower-level PPO agent, thereby affecting the scheduling of the entire environment state. Using target preference values ​​can better coordinate conflicts between multiple targets, and will not make the lower-level action selection optimize only target 1 or target 2, but will be biased on the basis of balance; the low-level PPO agent receives the optimization preference value of the high-level DQN agent and uses it as the target parameter of the reward function, and then selects the current scheduling point according to the action space scheduling algorithm rules based on the input environment state information. The action is right; The dual-agent real-time control workshop production process is as follows: (1) High-level DQN Target value in the DQN algorithm The calculation method is as follows: ; in, Is in state Next action After receiving the instant reward, is the discount factor, are the target network parameters, The target network For the next state Maximum Q value estimation; The error calculation formula between the estimated value and the target value in the current state, that is, the loss function, is as follows: ; The agent is in state Next action The immediate reward obtained from the environment reflects the direct effect of the current action, is the discount factor, is the next state of the target network The maximum Q value estimate, Is the main network's current state and actions Q value estimation; (2) Low-level PPO The following formula is the objective function of the PPO algorithm, by finding the parameters θ Maximize this function: ; It is The objective function of PPO at the iteration is: yes t The state of the moment, yes t Actions taken at all times, In the current state Take action strategy, is the importance sampling ratio, which represents the probability ratio of the new strategy to the old strategy, and is used to measure the change of the new strategy compared to the old strategy. Indicates limiting the importance sampling ratio to between, is a parameter that limits the range; It is The strategy parameters at the iteration, is the advantage function, It is The advantage function at the iteration, To evaluate the status Take action the pros and cons of; It means to obtain the minimum value among a set of target values; The decision-making part of the algorithm uses the Actor-Critic algorithm, which is divided into a Critic network and an Actor Policy Network. The former inputs the environment state and action to evaluate the value of taking the action under the current state. The latter inputs the current state and outputs the probability distribution of the action. The Critic network then uses it to evaluate the quality of the action and adjust the policy accordingly. The Critic network uses the mean square error loss function, that is, ; is the total number of samples, Corresponding to samples, is the mean square error loss function value obtained by the Critic network, and Calculate the cumulative reward value for each sample Comparing the status with the Critic network and actions Q-value estimation The mean square error between them promotes the Critic network to quickly correct the error, and finally The sum of the errors of the samples is averaged to obtain ; The optimization goals of the Actor network are as follows: ; is the policy function, Is the current state Take action The logarithmic probability of the policy function is used to measure the current policy in the state Next select action possibility; Is the advantage function advantage estimate, indicating that in this state Take action Compared with the average strategy, if it is greater than 0, it means it is better, then the probability of selecting the action is increased; if it is less than 0, it means it is worse, then the probability of selecting the action is reduced. The error of each sample is summed and averaged to get the loss function value of the Actor network .

4. The green dynamic multi-objective scheduling method for a flexible assembly workshop under the personnel learning effect according to claim 2 is characterized in that: Step S3, constructing a state space and action space that match the problem model and algorithm framework, and proposing an immediate reward function and a delayed reward function; (1) State space The production system for low-carbon flexible assembly workshop scheduling consists of three elements: workpiece, processing machine, and personnel. The global state characteristics include the three elements of workpiece characteristics, processing machine characteristics, and personnel characteristics. The workpiece characteristic indicators include actual completion time, remaining processing time, expected completion time and average completion time, as well as the predecessor-successor relationship between workpieces to indicate assembly operations; the processing machine characteristic indicators include machine utilization and machine completion time; The personnel characteristic indicators include workers’ initial learning rate, learning rate, and forgetting curve; Rescheduling point or decision point Average utilization of all devices on the , is the total number of devices: ; in, Indicates that at a rescheduling point or decision point On the device Equipment utilization rate; ; Representation device Completion time of the last process; At a rescheduling point or decision point The standard deviation of equipment utilization : ; All artifacts at rescheduling points or decision points Average completion rate : ; Indicates that at a rescheduling point or decision point Loading workpiece The number of completed processes, Indicates the Number of processes per workpiece; All processes are at the rescheduling point or decision point Average completion rate , is the total number of artifacts: ; in, Indicates that at a rescheduling point or decision point Loading workpiece completion rate; ; At a rescheduling point or decision point Estimated maximum completion time for loading workpieces : ; Indicates that at a rescheduling point or decision point Loading workpiece Completed process set, Indicates the process The running time, Indicates the process On the device 0-1 decision variables for the above operations; Workpiece Actual running time of completed operations By personnel learning rate Decide: ; Personnel learning rate The number of times personnel learn And the forgetting curve: ; Rescheduling point or decision point Energy consumption indicators for completing the process : ; in, Indicates that at a rescheduling point or decision point The actual energy consumption of the completed process, Indicates that at a rescheduling point or decision point The minimum energy consumption required to complete the process, Indicates that at a rescheduling point or decision point The median energy consumption required to complete the process; ; ; ; in, Indicates processing steps The operating energy consumption generated by Indicates the process Start time to the same device Previous process Idle energy consumption at the end time, Indicates the process used to execute Processing equipment set, Indicates that at a rescheduling point or decision point The maximum energy consumption required to complete the process, Indicates the process used to execute Processing personnel set, Indicates the process By workers On the device End time on; ; ; (2) Action Space The high-level DQN agent outputs the target emphasis probability ratio based on the current state, and the output result indicates the preference degree for the two targets in the current state; The low-level PPO agent generates the optimal process-machine action pair by combining the high-level target ratio and the current environment state. It constructs a heterogeneous graph neural network (HGNN), uses a multi-layer perceptron (MLP) to embed operation nodes and machine nodes, and builds a graph attention network (GAT) model to learn using the attention mechanism. It then filters out operation-machine node pairs that do not meet the scheduling conditions. Finally, it inputs the Actor network to obtain the action probability vector. When selecting action a, it uses ε-Greedy to add randomness to the exploration: ; is the strategy function of the Actor network, is a deterministic strategy, which means choosing The biggest move, It is a random strategy, calculated based on the Actor network Probability distribution randomly selects actions; (3) Reward Function The reward function includes immediate rewards and delayed rewards; Instant Rewards : Instant rewards include economic indicators and energy consumption indicators; Calculating economic indicators , based on the average equipment utilization , standard deviation of equipment utilization , the expected maximum completion time during training To calculate, Indicates the current time point, Indicates the next rescheduling point: Calculate energy consumption indicators , according to the energy consumption index , the minimum total energy consumption during training and current total energy consumption To calculate: ; Delayed Rewards : The delay reward is a negative value; the total completion time is generated after the lower-level agent selects the action pair and completes the scheduling process in the workshop environment and total energy consumption ,parameter and are the sum of the weighted probabilities of the two objectives in a complete scheduling process; 。

Citation Information

Patent Citations

  • Dynamic multi-target scheduling method for low-carbon distributed flexible job shop

    CN116500994A

  • Flexible workshop operation dynamic scheduling method based on deep reinforcement learning

    CN117892969A