A Deep Reinforcement Learning-Based Job Shop Scheduling Method Based on Multiple Bidding by Client Agents
By employing a deep reinforcement learning approach based on multiple bids from client agents, and combining deep reinforcement learning with a bidding mechanism, the conflict of personalized objectives in multi-agent job shop scheduling is resolved, generating a scheduling scheme acceptable to all parties, and achieving fair allocation and efficient utilization of resources.
Patent Information
- Application Number
- CN202510221851.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-02-27
AI Technical Summary
Existing technologies cannot effectively solve the scheduling problem of multi-agent job shops, especially the contradiction between the personalized target needs of customer agents and resource optimization, which makes traditional centralized optimization methods no longer applicable.
A deep reinforcement learning approach based on multiple bidding by client agents is adopted. Through a bidding mechanism, client agent agents bid for multiple processing steps, and workshop agent agents execute target setting decisions. By combining deep reinforcement learning and proximal policy optimization algorithms, personalized target scheduling of multiple agents is achieved.
It achieves the generation of scheduling schemes acceptable to all parties while protecting the private preferences of client agents, ensuring fair allocation and effective utilization of resources, and improving computational efficiency and solution performance.
Smart Images

Figure CN120070018B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of job shop scheduling technology, specifically relating to a deep reinforcement learning job shop scheduling method based on multiple bids by client agents. Background Technology
[0002] Job shop scheduling is the core of production management in discrete manufacturing enterprises, ensuring the orderly operation of the workshop. With the rapid development of the market economy, enterprise production has shifted from large-scale mass production to user-demand-driven personalized customization. This has led to a transformation in job shop scheduling from an optimization goal primarily focused on optimizing workshop resources to a personalized goal that considers and meets the differentiated needs of users while also optimizing workshop resources. Consequently, the traditional job shop scheduling problem has been transformed into a multi-agent job shop scheduling problem. The multi-agent job shop scheduling problem involves modeling customers as customer agents and jobs as job shop agents, meaning multiple agents collaborate to complete the job shop scheduling problem. The main research content of the job shop scheduling problem is how to effectively allocate workpieces from multiple competing customer agents on the limited machine resources owned by job shop agents to form a scheduling scheme acceptable to all parties.
[0003] The paper "Dong Z, Ren T, Qi F, et al. A reinforcement learning-based approach for solving multi-agent job shop scheduling problem. International Journal of Production Research, 2024: 1-26." discloses a multi-agent job shop scheduling optimization method based on deep reinforcement learning. This method proposes a graph Transformer network structure that combines graph neural networks and the Transformer model, which can quickly generate high-quality scheduling schemes. However, this graph Transformer network structure requires the target information of all clients to become the scheduling objective during the scheduling scheme generation process. Based on this scheduling objective, it solves the job shop scheduling problem with multiple competing client agents in a centralized optimization manner. All decisions in this scheduling optimization method are made by a central decision-maker, focusing on optimizing the overall system objective. Therefore, the scheduling scheme obtained through this graph Transformer network structure is difficult to meet the personalized objective requirements of multiple client agents.
[0004] Furthermore, due to reasons such as business competition or confidentiality, agents are usually unwilling to disclose their own target preferences, which makes the scheduling target based on all customer target information become the agent's private information. Consequently, the centralized optimization solution method for job shop scheduling is no longer applicable to such problems. Summary of the Invention
[0005] The purpose of this invention is to solve the problem of the inability to schedule multi-agent job shops in the existing technology, and to provide a deep reinforcement learning job shop scheduling method based on multiple bids by client agents. The bidding mechanism serves as an effective resource allocation framework. In each round of bidding, client agent agents trained by deep reinforcement learning bid for multiple processing steps. The job shop agent agents execute the bidding decision. Furthermore, while protecting the private preferences of each client agent, multiple agents with personalized goals participate in the scheduling decision-making to form a scheduling scheme acceptable to all parties.
[0006] To achieve the above objectives, the technical solution provided by this invention is:
[0007] A deep reinforcement learning-based job shop scheduling method based on multiple bids from client agents includes:
[0008] Step 1: Model a multi-agent job shop scheduling problem based on client agents and job shop agents, and obtain the client agent agent corresponding to the client agent, the job shop agent agent corresponding to the job shop, and the scheduling status of the job shop according to the multi-agent job shop scheduling problem;
[0009] Step 2: Based on the scheduling state, a feature extraction network is used to obtain the feature representation of each process node, a bidding strategy network is used to output the set of processes participating in the bidding in each round of bidding, and a near-end strategy optimization algorithm is used to train the client agent agent.
[0010] Step 3: Based on the feature representations of each process node obtained by the feature extraction network, the process scheduling order in each round of bidding is output by the calibration strategy network as the calibration decision order, and the workshop agent is trained using the strategy gradient algorithm.
[0011] Step 4: Jointly train the client agent intelligence trained in Step 2 and the workshop agent intelligence trained in Step 3 using a multi-agent deep deterministic strategy gradient algorithm. In the joint training algorithm, each client agent intelligence only observes its own process information to bid, while the workshop agent intelligence collects the bidding process sets of all client agents to make bidding decisions. The above bidding process is repeated continuously to generate the final scheduling scheme.
[0012] As a further limitation of the present invention, step one includes:
[0013] Step 11: Establish communication between the client agent and the job shop agent, and determine the set of agents participating in the scheduling decision; specifically, the set of agents participating in the scheduling decision is represented as follows:
[0014] {JSA,CA1,CA2,…,CA i ,…,CA N}Formula (1)
[0015] In formula (1), JSA represents the workshop agent, CA1 represents the customer agent for the first customer, CA2 represents the customer agent for the second customer, and CA... i CA represents the customer agent of the i-th customer. N This represents the customer agent for the Nth customer, where N represents the number of customer agents.
[0016] Step 12: Determine the set of machines for the job shop agency; specifically, the set of machines for the job shop agency is represented as follows:
[0017] M = {M1, M2, ..., M} m} Formula (2)
[0018] In formula (2), M represents the set of machines in the workshop that act as agents for JSA, M1 represents the first machine, M2 represents the second machine, and M... m This represents the m-th machine, where m represents the number of machines.
[0019] Step 13: Obtain the set of workpieces to be processed for each customer agent; specifically, the customer agent CA of the i-th customer... i The set of workpieces to be processed is represented as:
[0020]
[0021] In formula (3), This represents the customer agent CA of the i-th customer. i The set of workpieces to be processed This represents the customer agent CA of the i-th customer. i The first workpiece to be processed. This represents the customer agent CA of the i-th customer. i The second workpiece to be processed. This represents the customer agent CA of the i-th customer. i The j-th workpiece to be processed This represents the customer agent CA of the i-th customer. i The nth i There are 1 workpiece to be processed; among which:
[0022] (1) The customer agent CA of the i-th customer i The j-th workpiece to be processed The completion time is Delivery period is
[0023]
[0024] (2) The customer agent CA of the i-th customer i The set of workpieces to be processed In the diagram, the process information for each workpiece to be processed is represented as follows:
[0025]
[0026] In formula (4), This represents the customer agent CA of the i-th customer. i The j-th workpiece to be processed The kth process In the machine The processing time is as follows: Where, j∈[1,n] i ], k∈[1,m];
[0027] Step 14: Determine the objectives of the client agent; specifically, the client agent CA of the i-th client. i target f i Includes: minimizing the maximum completion time C i Minimize the total completion time (TC) i and minimize total delay TT i , is represented as:
[0028]
[0029] In formula (5), Indicates the completion time, n i Indicates a total of n i One workpiece to be processed. Indicates the delivery date;
[0030] Step 15: Determine the objective of the workshop agent to minimize the maximum completion time; specifically, minimize the maximum completion time C. max The calculation expression is:
[0031]
[0032] In formula (6), i represents the i-th customer, C i This represents minimizing the maximum completion time;
[0033] Step 16: Model the client agent and the job shop agent separately to obtain the client agent agent and the job shop agent agent; specifically, the client agent agent is represented as CA. i The workshop agent is represented by JSA; where the client agent is CA. i Bidding decisions are made based on the bidding strategy network, and the workshop agent JSA executes the bid awarding decisions based on the bid awarding strategy network.
[0034] Step 17: Construct the job shop scheduling environment and obtain the scheduling status of the job shops; specifically, based on the job shop scheduling environment, use a disjunctive graph model to describe the scheduling status of the job shops; wherein, the disjunctive graph model is represented as:
[0035]
[0036] In formula (7), Represents the disjunctive graph model. Represents a set of process nodes. This represents the set of directed arcs representing adjacent processes on the same workpiece. This represents the set of undirected arcs representing adjacent processes on the same machine.
[0037] As a further limitation of the present invention, step two includes:
[0038] Step 21: Initialize the client agent intelligent agent; specifically, the client agent intelligent agent includes a feature extraction network, a bidding strategy network, and a state value network; the feature extraction network includes a graph neural network and an encoder network; the bidding strategy network and the state value network both adopt a fully connected neural network structure;
[0039] Step 22: The feature extraction network in the client agent intelligent agent obtains the feature representation of each process node; specifically, firstly, the graph neural network in the feature extraction network is used to extract the feature representation of each process node in the disjunctive graph model; then, the feature representation of each process node is input into the encoder network in the feature extraction network, and the final feature representation of each process node is obtained based on the encoder network.
[0040] Step 23: Input the feature representations of all process nodes obtained in Step 22 into the bidding strategy network, and output the probability of selecting each process node based on the bidding strategy network; specifically, select m currently processable processes into the bidding process set according to the probability values in descending order, denoted as... If the number of currently processable operations is less than m, then all processable operations will be added to the bidding operation set.
[0041] Step 24: Input the feature representations of all processable process nodes obtained in Step 22 into the state value network, and obtain the estimated values of all current processable processes based on the state value network.
[0042] Step 25: Perform the bidding process set obtained in step 23. Update the job shop scheduling environment status. t ;
[0043] Step 26: Obtain the target type based on the client agent's objectives, and calculate the reward value based on the target type. Specifically, the reward value is calculated using the following formula.
[0044]
[0045] In formula (8), This represents the customer agent CA corresponding to the i-th customer. i Reward value f i (s t ) represents state s t Customer Agent CA i The objective function value; f i (s t+1 ) represents state s t+1 Customer Agent CA i The objective function value;
[0046] Step 27: Repeat steps 22-26 until all processes are scheduled.
[0047] Step 28: Train the client agent using the proximal policy optimization algorithm;
[0048] Step 29: Repeat steps 22-28 until the client agent policy converges.
[0049] As a further limitation of the present invention, step 22 includes:
[0050] Step 221: The features of the process nodes in the disjunctive graph model include: the estimated earliest completion time of the process node. Has the scheduling been completed?
[0051] Step 222: The graph neural network performs K iterations to extract the embedding vectors of the process nodes. The formula for calculating the qth iteration is:
[0052]
[0053] In formula (9), Indicates process node The embedding vector in the q-th iteration, MLP (q) Let represent the q-th fully connected layer of a graph neural network, where ∈ denotes a learnable parameter. Indicates process node The original characteristics, Indicates process node The set of neighboring nodes, Let represent the embedding vector of the neighboring node u in the (q-1)th iteration;
[0054] Step 223: Extract the embedding vector from the graph neural network and perform average pooling on it to obtain the graph embedding vector. Among them, graph embedding vector The calculation formula is:
[0055]
[0056] In formula (10), n i Indicates a total of n i There are 10 workpieces to be processed, and m represents the number of machines. Indicates process node, Represents a set of process nodes. Indicates process node Node embedding vectors;
[0057] Step 224: Insert the node embedding vector obtained in step 222. Input the encoder network, and obtain the enhanced node embedding vector based on the encoder network, expressed as:
[0058]
[0059] In formula (11), Indicates process node The set of neighboring nodes, This represents the final augmented node embedding vector. Indicates attention weights, This represents a linear mapping of the node embedding vector;
[0060] Step 225: Combine the enhanced node embedding vector obtained in step 224 with the graph embedding vector obtained in step 223. The final node features obtained by concatenating the individual nodes are represented as follows:
[0061]
[0062] In formula (12), Represents the final node characteristics, Represents the augmented node embedding vector. Represents a graph embedding vector.
[0063] As a further limitation of the present invention, step three includes:
[0064] Step 31: Initialize the job shop agent; specifically, the job shop agent includes a feature extraction network and a calibration policy network;
[0065] Step 32: Based on the client agent intelligence in Step 2, obtain the feature representation of each process node in each client agent using Step 22;
[0066] Step 33: Input the feature representations of all process nodes obtained in Step 32 into the calibration strategy network of the workshop agent, and the calibration strategy network, as a decoder, outputs the calibration decision sequence in an autoregressive manner;
[0067] Step 34: Based on the calibration decision sequence obtained in step 33, execute m calibration schemes respectively, and calculate the reward value corresponding to each of the m calibration schemes;
[0068] Step 35: After iterating through steps 32-34 B times, the workshop agent is trained using the policy gradient algorithm; specifically, the calculation expression for the policy gradient algorithm is:
[0069]
[0070] In formula (13), Let B represent the gradient of the objective function; B represents the number of iterations; and m represents the number of machines. Representation scheme The reward value, μ(τ) b () represents the mean reward value of all solutions in the b-th iteration. σ(τ b ) represents the standard deviation of the reward values of all solutions in the b-th iteration. ent coeff The coefficient representing entropy, and entropy representing the entropy value;
[0071] Step 36: Repeat steps 32-35 until the job shop agent strategy converges.
[0072] As a further limitation of the present invention, step four includes:
[0073] Step 41: Load the client agent intelligent agent trained in Step 2, and load the job shop agent intelligent agent trained in Step 3; specifically, load the parameters of the feature extraction network, bidding strategy network, and state value network in the client agent intelligent agent; load the parameters of the calibration strategy network in the job shop agent intelligent agent;
[0074] Step 42: Using the client agent intelligent agent from step two, extract the feature representation of each process node for each client agent's respective client agent intelligent agent;
[0075] Step 43: Each client agent's bidding strategy network takes the final node features of each process in the set of processes to be processed at the current decision moment as input, and outputs its own set of bidding processes, thus obtaining the bid set of all client agent agents. set Specifically, by aggregating the bidding process sets of each of the N client agent intelligent agents, we obtain the bidding process set (bid) for all client agent intelligent agents. set This can be represented as:
[0076] Step 44: The workshop agent receives the bid set of all customer agents. set The system calculates the final node characteristics of each process and uses the calibration strategy network to perform autoregressive calibration decisions to obtain calibration decision results. During the execution of the autoregressive calibration decision, when the length of the calibration sequence corresponding to the autoregressive calibration decision reaches the number of machines m or all bidding processes have been scheduled, the system stops executing the output of the autoregressive calibration decision on the calibration strategy network.
[0077] Step 45: Execute the calibration decision results obtained in step 44, and update the multi-agent job shop scheduling environment based on the job shop agent agent.
[0078] Step 46: Calculate the reward value for each client agent based on their target type. The calculation formula is as follows:
[0079]
[0080] In formula (14), Indicates the reward value. This represents the customer agent CA of the i-th customer. i The number of scheduled processes, Indicates the scheduled process nodes The value of the state value network output by the client agent is the value of the scheduled process nodes.
[0081] Step 47: Based on the value of the process node, calculate the reward value of the multi-agent workshop. The calculation formula is as follows:
[0082]
[0083] In formula (15), This represents the reward value for multi-agent workshops, where N represents the number of customer agents. This represents the customer agent CA of the i-th customer. i The number of scheduled processes, Indicates the scheduled process nodes Value;
[0084] Step 48: Repeat steps 42-47 until all processes of each customer agent are scheduled.
[0085] Step 49: Train the client agent and the job shop agent using the multi-agent deep deterministic policy gradient algorithm;
[0086] Step 410: Repeat steps 42-49 until the strategies of the client agent and the job shop agent converge.
[0087] The advantages of this invention are:
[0088] 1. The multi-agent job shop scheduling method of the present invention is based on deep reinforcement learning and bidding mechanism to schedule all processes of the client agent. Compared with the traditional scheduling method, the present invention not only allows the client agent and the resource agent of the job shop to participate in the scheduling decision, but also allows the client agent agent trained by deep reinforcement learning to bid for multiple processes to be processed in each round of bidding, and the job shop agent agent to execute the target decision, so as to realize the joint participation of multiple agents with personalized goals in the scheduling decision.
[0089] 2. In the scheduling decision-making process, this invention takes the personalized needs of the client agents as the core consideration, so that the scheduling scheme generated by this invention fully reflects the preference information of each client agent; this invention adopts the method of each client agent bidding multiple times in each round of bidding, which greatly satisfies the personalized needs of client agents to fully express their private preferences.
[0090] 3. The job shop scheduling method of the present invention combines deep reinforcement learning with a bidding mechanism. The bidding mechanism ensures the fair allocation of job shop resources among various client agents. Deep reinforcement learning enables client agents and job shop agents to autonomously learn resource conflict negotiation methods in multiple rounds of bidding, ensuring the effective utilization of job shop resources. In addition, the client agent agent and job shop agent agent of the present invention are trained by a joint training algorithm, which has good computational efficiency and solution performance.
[0091] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0092] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:
[0093] Figure 1 This invention provides a flowchart of a deep reinforcement learning job shop scheduling method based on multiple bids by client agents;
[0094] Figure 2 : A schematic diagram of the client agent intelligent agent bidding algorithm provided by this invention;
[0095] Figure 3 : A schematic diagram of the calibration algorithm for the agent intelligence in the workshop provided by this invention;
[0096] Figure 4 This invention provides a schematic diagram of a joint training algorithm for client agent intelligence and job shop agent intelligence. Detailed Implementation
[0097] The embodiments of the present invention are described in detail below. These embodiments are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.
[0098] Please see Figure 1 This invention provides a deep reinforcement learning-based job shop scheduling method based on multiple bids from client agents, comprising:
[0099] Step 1: Model the multi-agent job shop scheduling problem based on client agents and job shop agents, and obtain the client agent agent corresponding to the client agent, the job shop agent agent corresponding to the job shop, and the scheduling status of the job shop according to the multi-agent job shop scheduling problem.
[0100] Step one of this embodiment of the invention includes:
[0101] Step 11: Establish communication between the client agent and the job shop agent, and determine the set of agents participating in the scheduling decision; specifically, the set of agents participating in the scheduling decision is represented as follows:
[0102] {JSA,CA1,CA2,…,CA i ,…,CA N}Formula (1)
[0103] In formula (1), JSA represents the workshop agent, CA1 represents the customer agent for the first customer, CA2 represents the customer agent for the second customer, and CA... i CA represents the customer agent of the i-th customer. N This represents the customer agent for the Nth customer, where N represents the number of customer agents.
[0104] Step 12: Determine the set of machines for the job shop agency; specifically, the set of machines for the job shop agency is represented as follows:
[0105] M = {M1, M2, ..., M} m} Formula (2)
[0106] In formula (2), M represents the set of machines in the workshop that act as agents for JSA, M1 represents the first machine, M2 represents the second machine, and M... m This represents the m-th machine, where m represents the number of machines.
[0107] Step 13: Obtain the set of workpieces to be processed for each customer agent; specifically, the customer agent CA of the i-th customer... i The set of workpieces to be processed is represented as:
[0108]
[0109] In formula (3), This represents the customer agent CA of the i-th customer. i The set of workpieces to be processed This represents the customer agent CA of the i-th customer. i The first workpiece to be processed. This represents the customer agent CA of the i-th customer. i The second workpiece to be processed. This represents the customer agent CA of the i-th customer. i The j-th workpiece to be processed This represents the customer agent CA of the i-th customer. i The nth i There are 1 workpiece to be processed; among which:
[0110] (1) The customer agent CA of the i-th customer i The j-th workpiece to be processed The completion time is Delivery period is
[0111]
[0112] (2) The customer agent CA of the i-th customer i The set of workpieces to be processed In the diagram, the process information for each workpiece to be processed is represented as follows:
[0113]
[0114] In formula (4), This represents the customer agent CA of the i-th customer. i The j-th workpiece to be processed The kth process In the machine The processing time is as follows: Where, j∈[1,n] i ], k∈[1,m];
[0115] Step 14: Determine the objectives of the client agent; specifically, the client agent CA of the i-th client. i target f i Includes: minimizing the maximum completion time C i Minimize the total completion time (TC) i and minimize total delay TT i , is represented as:
[0116]
[0117] In formula (5), Indicates the completion time, n i Indicates a total of n i One workpiece to be processed. Indicates the delivery date;
[0118] Step 15: Determine the objective of the workshop agent to minimize the maximum completion time; specifically, minimize the maximum completion time C. max The calculation expression is:
[0119]
[0120] In formula (6), i represents the i-th customer, C i This represents minimizing the maximum completion time;
[0121] Step 16: Model the client agent and the job shop agent separately to obtain the client agent agent and the job shop agent agent; specifically, the client agent agent is represented as CA. iThe workshop agent is represented by JSA; where the client agent is CA. i Bidding decisions are made based on the bidding strategy network, and the workshop agent JSA executes the bid awarding decisions based on the bid awarding strategy network.
[0122] Step 17: Construct the job shop scheduling environment and obtain the scheduling status of the job shops; specifically, based on the job shop scheduling environment, use a disjunctive graph model to describe the scheduling status of the job shops; wherein, the disjunctive graph model is represented as:
[0123]
[0124] In formula (7), Represents the disjunctive graph model. Represents a set of process nodes. This represents the set of directed arcs representing adjacent processes on the same workpiece. This represents the set of undirected arcs representing adjacent processes on the same machine.
[0125] Step 2: Based on the scheduling state, a feature extraction network is used to obtain the feature representation of each process node, a bidding strategy network is used to output the set of processes participating in the bidding in each round of bidding, and a near-end strategy optimization algorithm is used to train the client agent agent.
[0126] Step two of this embodiment of the invention includes:
[0127] Step 21: Initialize the client agent intelligent agent; specifically, the client agent intelligent agent includes a feature extraction network, a bidding strategy network, and a state value network; the feature extraction network includes a graph neural network and an encoder network; both the bidding strategy network and the state value network adopt a fully connected neural network structure;
[0128] Step 22: The feature extraction network in the client agent intelligent agent obtains the feature representation of each process node; specifically, firstly, the graph neural network in the feature extraction network is used to extract the feature representation of each process node in the disjunctive graph model; then, the feature representation of each process node is input into the encoder network in the feature extraction network, and the final feature representation of each process node is obtained based on the encoder network.
[0129] Step 23: Input the feature representations of all process nodes obtained in Step 22 into the bidding strategy network, and output the probability of selecting each process node based on the bidding strategy network; specifically, select m currently processable processes into the bidding process set according to the probability values in descending order, denoted as... If the number of currently processable operations is less than m, then all processable operations will be added to the bidding operation set.
[0130] Step 24: Input the feature representations of all processable nodes obtained in Step 22 into the state value network, and obtain the estimated values of all current processable nodes based on the state value network.
[0131] Step 25: Execute the bidding process set obtained in Step 23. Update the job shop scheduling environment status. t ;
[0132] Step 26: Obtain the target type based on the client agent's goals, and calculate the reward value based on the target type. Specifically, the reward value is calculated using the following formula.
[0133]
[0134] In formula (8), This represents the customer agent CA corresponding to the i-th customer. i Reward value f i (s t ) represents state s t Customer Agent CA i The objective function value; f i (s t+1 ) represents state s t+1 Customer Agent CA i The objective function value;
[0135] Step 27: Repeat steps 22-26 until all processes are scheduled.
[0136] Step 28: Train the client agent using the proximal policy optimization algorithm;
[0137] Step 29: Repeat steps 22-28 until the client agent policy converges.
[0138] More specifically, step 22 in this embodiment of the invention includes:
[0139] Step 221: The characteristics of the process nodes in the disjunctive graph model include: the estimated earliest completion time of the process node. Has the scheduling been completed?
[0140] Step 222: The graph neural network performs K iterations to extract the embedding vectors of the process nodes. The formula for calculating the qth iteration is:
[0141]
[0142] In formula (9), Indicates process node The embedding vector in the q-th iteration, MLP (q) Let represent the q-th fully connected layer of a graph neural network, where ∈ denotes a learnable parameter. Indicates process node The original characteristics, Indicates process node The set of neighboring nodes, Let represent the embedding vector of the neighboring node u in the (q-1)th iteration;
[0143] Step 223: Extract the embedding vector from the graph neural network and perform average pooling on it to obtain the graph embedding vector. Among them, graph embedding vector The calculation formula is:
[0144]
[0145] In formula (10), n i Indicates a total of n i There are 10 workpieces to be processed, and m represents the number of machines. Indicates process node, Represents a set of process nodes. Indicates process node Node embedding vectors;
[0146] Step 224: Insert the node embedding vector obtained in step 222. Input the encoder network, and obtain the augmented node embedding vectors based on the encoder network, expressed as:
[0147]
[0148] In formula (11), Indicates process node The set of neighboring nodes, This represents the final augmented node embedding vector. Indicates attention weights, This represents a linear mapping of the node embedding vector;
[0149] Step 225: Combine the enhanced node embedding vector obtained in step 224 with the graph embedding vector obtained in step 223. The final node features obtained by concatenating the individual nodes are represented as follows:
[0150]
[0151] In formula (12), Represents the final node characteristics, Represents the augmented node embedding vector. Represents a graph embedding vector.
[0152] Please see Figure 2 In practical applications, the client agent intelligent agent described in this embodiment of the invention first initializes the job shop scheduling environment and the client agent intelligent agent's feature extraction network, bidding strategy network, and state value network. The feature extraction network includes a graph neural network and an encoder network. The feature extraction network is used to obtain the feature representation of each process node. Specifically, the feature extraction network uses a graph neural network to obtain the embedding vector of each process node, and performs average pooling on the node embedding vector to obtain a graph embedding vector. The node embedding vector is input into the encoder network to obtain the encoder embedding vector, and the graph embedding vector and the encoder embedding vector are concatenated to obtain the final node feature. The final node feature of the current processable process is input into the bidding strategy network to output the probability of selecting each process. At most m current processable processes are selected into the bidding process set according to the probability value from largest to smallest. The final node feature of the current processable process is input into the state value network to obtain its state value. The bidding process set is executed, and the job shop scheduling environment is updated. The above operations are repeated until all processes are scheduled. After all processes of the client agent are scheduled, the client agent intelligent agent is trained using a proximal policy optimization algorithm.
[0153] Step 3: Based on the feature representations of each process node obtained by the feature extraction network, the process scheduling order in each round of bidding is output by the calibration strategy network as the calibration decision order, and the workshop agent is trained using the policy gradient algorithm.
[0154] Specifically, step three in this embodiment of the invention includes:
[0155] Step 31: Initialize the job shop agent; specifically, the job shop agent includes a feature extraction network and a calibration policy network;
[0156] Step 32: Based on the client agent intelligent agent in Step 2, use Step 22 to obtain the feature representation of each process node in each client agent;
[0157] Step 33: Input the feature representations of all process nodes obtained in Step 32 into the calibration strategy network of the work shop agent. The calibration strategy network acts as a decoder and outputs the calibration decision sequence in an autoregressive manner.
[0158] Step 34: Based on the calibration decision sequence obtained in Step 33, execute m calibration schemes respectively, and calculate the reward value corresponding to each of the m calibration schemes;
[0159] Step 35: After iterating through steps 32-34 B times, train the job shop agent using the policy gradient algorithm; specifically, the calculation expression for the policy gradient algorithm is:
[0160]
[0161] In formula (13), Let B represent the gradient of the objective function; B represents the number of iterations; and m represents the number of machines. Representation scheme The reward value, μ(τ) b () represents the mean reward value of all solutions in the b-th iteration. σ(τ b ) represents the standard deviation of the reward values of all solutions in the b-th iteration. ent coeff The coefficient representing entropy, and entropy representing the entropy value;
[0162] Step 36: Repeat steps 32-35 until the job shop agent strategy converges.
[0163] Initialize the job shop scheduling environment and the feature extraction network and scaling policy network of the job shop agent. The encoder embedding vector and graph embedding vector obtained by the feature extraction network are concatenated to form the final node features. The final node features of each node are input as key and value vectors into the multi-pointer network. The final node features of each node, together with the multiple scheduling sequences, constitute the augmented state embedding vector, which is input as the query vector into the multi-pointer network. The multi-pointer network receives the query vector, key vector, and value vector, and its output is passed through a linear layer and a softmax layer to obtain the probability of selecting each processing step. Based on this probability, m scheduling sequences are generated to balance the mean and variance during training. The m scheduling sequences are executed respectively, and the augmented state embedding vector is updated. This process is repeated until all processes of the job shop agent are scheduled. After all processes are scheduled, the job shop agent agent is trained using the policy gradient algorithm.
[0164] Step 4: Jointly train the client agent agent trained in Step 2 and the job shop agent agent trained in Step 3 using a multi-agent deep deterministic strategy gradient algorithm. In the joint training algorithm, each client agent agent only observes its own process information and bids multiple times, while the job shop agent agent collects the bidding process sets of all client agents to make bidding decisions. The above bidding process is repeated continuously to generate the final scheduling scheme.
[0165] Specifically, step four of the present invention includes:
[0166] Step 41: Load the client agent intelligent agent trained in Step 2, and load the job shop agent intelligent agent trained in Step 3; specifically, load the parameters of the feature extraction network, bidding strategy network, and state value network in the client agent intelligent agent; load the parameters of the calibration strategy network in the job shop agent intelligent agent.
[0167] Step 42: Using the client agent intelligent agent from Step 2, extract the feature representation of each process node for each client agent's respective client agent intelligent agent;
[0168] Step 43: Each client agent's bidding strategy network uses the final node features of each process in the set of processes to be processed at the current decision moment as input. The bidding strategy network outputs its own set of bidding processes, thus obtaining the bid set of all client agent agents. set Specifically, by aggregating the bidding process sets of each of the N client agent intelligent agents, we obtain the bidding process set (bid) for all client agent intelligent agents. set This can be represented as:
[0169] Step 44: The workshop agent receives the set of bids for work processes from all customer agents. set The system also considers the characteristics of the final nodes of each process and uses a calibration strategy network to perform autoregressive calibration decisions to obtain calibration decision results. During the execution of the autoregressive calibration decision, when the length of the calibration sequence corresponding to the autoregressive calibration decision reaches the number of machines m or all bidding processes have been scheduled, the system stops executing the output of the autoregressive calibration decision on the calibration strategy network.
[0170] Step 45: Execute the calibration decision results obtained in step 44, and update the multi-agent job shop scheduling environment based on the job shop agent agent.
[0171] Step 46: Calculate the reward value for each client agent based on their target type. The calculation formula is as follows:
[0172]
[0173] In formula (14), Indicates the reward value. This represents the customer agent CA of the i-th customer. i The number of scheduled processes, Indicates the scheduled process nodes The value of the client agent's state value network output is the value of the scheduled process nodes.
[0174] Step 47: Based on the value of the process node, calculate the reward value of the multi-agent workshop. The calculation formula is as follows:
[0175]
[0176] In formula (15), This represents the reward value for multi-agent workshops, where N represents the number of customer agents. This represents the customer agent CA of the i-th customer. i The number of scheduled processes, Indicates the scheduled process nodes Value;
[0177] Step 48: Repeat steps 42-47 until all processes of each customer agent are scheduled.
[0178] Step 49: Train the client agent and the job shop agent using the multi-agent deep deterministic policy gradient algorithm;
[0179] Step 410: Repeat steps 42-49 until the strategies of the client agent and the job shop agent converge.
[0180] Initialize the job shop scheduling environment and load the client agent agents trained in step two and the job shop agent agents trained in step three. The feature extraction network of the client agent agents obtains the final node features of each process. All client agent agents share a single graph neural network to improve computational efficiency. Each client agent agent's bidding strategy network takes the final node features of its respective process as input and outputs its own set of bidding processes. The job shop agent agents receive the sets of bidding processes and their corresponding final node features from all client agent agents and execute autoregressive calibration decisions using a calibration strategy network. During the calibration decision process, when the calibration sequence length reaches the number of machines or all bidding processes have been scheduled, the autoregressive output stops, and a scheduling sequence is obtained. The scheduling sequence is executed, and the job shop scheduling environment state is updated. This process is repeated until all client agent processes are scheduled. After all processes are scheduled, the client agent agents and job shop agent agents are trained using a multi-agent deep deterministic policy gradient algorithm.
[0181] It should be noted that in this embodiment of the invention, the client agent bids multiple times. In each round of bidding, the client agent can bid on multiple items, whereas in the case where only one item can be bid on per round, the agent in this embodiment is a modeling technique that simulates and solves real-world problems by abstracting entities in complex systems into individuals with autonomy, interactivity, and intelligence. Each agent has its own goals, behaviors, and decision-making capabilities, and can interact with other agents to jointly complete scheduling tasks. In this embodiment, the client agent is a common term in scheduling problems, referring to the use of agent technology to describe customer needs and the activities of customer needs interacting with the job shop.
[0182] This invention discloses a deep reinforcement learning-based job shop scheduling method based on multiple bidding by client agents. In each round of bidding, client agent agents trained by deep reinforcement learning bid for multiple processing steps. The job shop agent agents execute calibration decisions, enabling multiple agents with personalized goals to participate in the scheduling decision-making process. Specifically, at each decision moment, the bidding strategy network of the client agent agents takes its private workpiece information as input and outputs the processes to be bid for. The calibration strategy network of the job shop agent agents takes the bidding process information as input, outputs the calibration decision result, and submits the calibration decision result to the job shop scheduling environment to update the disjunctive graph model. This bidding process is repeated until all client agent processes are scheduled. This job shop scheduling method combines deep reinforcement learning and a bidding mechanism, effectively solving the problem of collaborative scheduling of multiple agents with personalized goals, and exhibits good computational efficiency and optimization effects. Furthermore, compared to traditional scheduling methods, this embodiment of the invention allows client agents and resource agents to jointly participate in scheduling decisions, and in the decision-making process, their own personalized needs are the sole consideration, so that the resulting scheduling scheme can fully reflect the preference information of each agent; the combination of deep reinforcement learning and bidding mechanism greatly improves resource allocation efficiency and has good computational efficiency and solution performance; in this embodiment of the invention, the client agent makes multiple bids in each round of bidding, which fully expresses the private preferences of the client agent.
[0183] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the scope of the technology disclosed in the present invention, and such modifications or substitutions should all be covered within the scope of protection of the present invention.
Claims
1. A deep reinforcement learning job-shop scheduling method based on customer-agent multi-bidding, characterized in that, The method comprises the following steps: Step one, modeling a multi-agent job shop scheduling problem based on a customer agent and a job shop agent, and obtaining a customer agent corresponding to the customer agent, a job shop agent corresponding to the job shop, and a scheduling state of the job shop by using a disjunctive graph model according to the multi-agent job shop scheduling problem; wherein the disjunctive graph in the disjunctive graph model takes a process as a node, and a non-directed arc in the disjunctive graph represents a scheduling sequence of adjacent processes on the same machine, and the scheduling state is updated by the sequence of process scheduling in the scheduling process; Step two, obtaining a feature representation of each process node by using a feature extraction network based on the scheduling state, outputting a set of processes participating in bidding in each round of bidding by using a bidding strategy network according to the feature representation of each process node, and training the customer agent by using a proximal policy optimization algorithm, so that the customer agent generates a set of bidding processes based only on the feature of its own process node; wherein the feature representation of the process node in the disjunctive graph model includes the estimated earliest completion time of the process node and the completion scheduling flag; Step three, obtaining the feature representation of each process node in step two by using the feature extraction network, outputting the process scheduling sequence in each round of bidding as a benchmark decision sequence in an autoregressive manner by using a benchmark strategy network, and training the job shop agent by using a policy gradient algorithm; Step four, jointly training the customer agent after step two and the job shop agent after step three by using a multi-agent deep deterministic policy gradient algorithm; wherein each customer agent in the joint training algorithm only observes its own process information for multi-bidding, and the job shop agent collects the bidding process set of all customer agents for autoregressive output of the benchmark decision, and the above bidding and benchmarking process is repeatedly repeated to generate a final scheduling scheme.
2. The deep reinforcement learning job shop scheduling method based on customer agent multi-bidding according to claim 1, wherein, The step one comprises: Step 11, establishing communication between the customer agent and the job shop agent, and determining a set of agents participating in scheduling decision; specifically, the set of agents participating in scheduling decision is represented as: { JS A, CA1, CA2,..., CA i ,..., CA N} Equation (1) In Equation (1), JSA denotes a job shop agent, CA1 denotes a customer agent of a first customer, CA2 denotes a customer agent of a second customer, CA i denotes a customer agent of an i-th customer, CA N denotes a customer agent of an N-th customer, and N denotes the number of customer agents. Step 12, determining a machine set of the job shop agent; specifically, the machine set of the job shop agent is represented as: M = {M1, M2,..., M m} Equation (2) In Equation (2), M represents a machine set of the job shop agent JSA, M1 represents a first machine, M2 represents a second machine, M m represents an mth machine, and m represents the number of machines. In Equation (2), M represents a machine set of the job shop agent JSA, M1 represents a first machine, M2 represents a second machine, M m represents an mth machine, and m represents the number of machines. Step 13, obtaining the set of workpieces to be processed by each customer agent; specifically, the set of workpieces to be processed by the customer agent CA i of the i-th customer is represented as: In formula (3), This represents the customer agent CA of the i-th customer. i The set of workpieces to be processed This represents the customer agent CA of the i-th customer. i The first workpiece to be processed. This represents the customer agent CA of the i-th customer. i The second workpiece to be processed. This represents the customer agent CA of the i-th customer. i The j-th workpiece to be processed This represents the customer agent CA of the i-th customer. i The nth i There are 1 workpiece to be processed; among which: (1) the client agent CA of the i-th client i in the j-th workpiece to be processed the completion time is the delivery time is (2) a client agent CA of the i-th client i a set of workpieces to be processed In the above, the process information of each workpiece to be processed is represented as: In formula (4), a client agent CA representing the i-th client i the j-th workpiece to be processed in the i-th client the k-th process on the machine with a processing time of wherein j∈[1,ni], k∈[1,m]; Step 14, determining the target of the customer agent; in particular, the target f of the customer agent CA of the i-th customer i Step 14, determining the target of the customer agent; in particular, the target f of the customer agent CA of the i-th customer i Step 14, determining the target of the customer agent; in particular, the target f of the customer agent CA of the i-th customer i Step 14, determining the target of the customer agent; in particular, the target f of the customer agent CA of the i-th customer i Step 14, determining the target of the customer agent; in particular, the target f of the customer agent CA of the i-th customer i Step 14, determining the In formula (5), denotes the completion time, n i denotes the total number of n i workpieces to be processed, denotes the delivery date; Step 15, determining a target minimum maximum completion time for the job shop agent; in particular, a minimum maximum completion time C max The computational expression is: C max = max i C i Formula (6) In Equation (6), i denotes the ith customer, C i denotes the minimization of the maximum completion time; Step 16, modeling the customer agent and the job shop agent respectively to obtain a customer agent and a job shop agent; specifically, the customer agent is represented as CA i , and the job shop agent is represented as JSA; wherein the customer agent CA i performs bidding decision according to the bidding strategy network, and the job shop agent JSA performs bidding decision according to the bidding strategy network. Step 17, constructing a job shop scheduling environment, and obtaining a scheduling state of the job shop; specifically, the scheduling state of the job shop is described by using a disjunctive graph model according to the job shop scheduling environment; wherein the disjunctive graph model is represented as: In formula (7), denotes a disjunctive graph model, denotes a set of process nodes, denotes a set of directed arcs between adjacent processes on the same workpiece, denotes a set of undirected arcs between adjacent processes on the same machine. 3.The deep reinforcement learning job-shop scheduling method based on customer-agent multi-bidder of claim 1, wherein, The step two comprises: Step 21, initializing the customer agent; specifically, the customer agent comprises a feature extraction network, a bidding strategy network and a state value network; the feature extraction network comprises a graph neural network and an encoder network; the bidding strategy network and the state value network both adopt a fully connected neural network structure; Step 22, the feature extraction network in the customer agent obtains a feature representation of each process node; specifically, first, the graph neural network in the feature extraction network is used to extract the feature representation of each process node in the disjunctive graph model; Then, the feature representation of each process node is input into the encoder network of the feature extraction network, and a final feature representation of each process node is obtained based on the encoder network; Step 23, input the feature representation of all process nodes obtained in the step 22 into the bidding strategy network, and output the probability of selecting each process node based on the bidding strategy network; specifically, according to the order from large to small of the probability value, select d current processable process into the bidding process set, denoted as denotes the dth process of the ith customer agent in the tth round of bidding; if the number of current processable processes is less than d, all processable processes are added to the bidding process set Step 24, inputting the feature representation of all processable process nodes obtained in the step 22 into the state value network, obtaining the estimated value of all processable processes based on the state value network wherein, represents any one of the current all processable process nodes; Step 25, executing the bidding procedure set obtained in said step 23 updating the job shop scheduling environment state s t ; Step 26, the target type is obtained based on the target of the client agent, and a reward value is calculated through the target type Specifically, the reward value is calculated through the following formula In formula (8), represents the reward value of the customer agent CA corresponding to the i-th customer i f i (s t ) represents the objective function value of the customer agent CA in the state s t i i (s t+1 ) represents the objective function value of the customer agent CA in the state s t+1 i ; Step 27, repeat steps 22-26 until all processes are completed. Step 28, train the customer agent using a proximal policy optimization algorithm; Step 29, repeat steps 22-28 until the customer agent policy converges.
4. The deep reinforcement learning job shop scheduling method based on customer agent multi-bidding according to claim 3, characterized in that, The step 22 includes: Step 221, the features of the process node in the extraction graph model include: the estimated earliest completion time of the process node Whether the scheduling is completed Step 222, the graph neural network performs K times of iteration to extract the embedding vector of the process node Process node in the qth round of graph neural network The embedding vector calculation formula of the process node is: In formula (9), denotes a process node denotes an embedding vector of the qth layer of the MLP (q) denotes a fully connected neural network of the qth layer of the graph neural network, ∈ (q) denotes a learnable parameter, denotes a process node denotes an original feature of the process node denotes a process node denotes a set of neighborhood nodes of the process node denotes an embedding vector of the q-1th layer of the neighborhood node u; Step 223, extract the embedding vector from the graph neural network, and average pool it to obtain a graph embedding vector wherein the graph embedding vector The calculation formula is: In formula (10), n i represents the total number of n i workpieces to be processed, m represents the number of machines, represents a process node, represents a set of process nodes, represents a node embedding vector of a process node . Step 224, embedding the node in the vector obtained in step 222 inputting the encoder network, obtaining an enhanced node embedding vector based on the encoder network, expressed as: In Equation (11), denotes a set of neighborhood nodes of the procedure node , denotes the resulting enhanced node embedding vector, denotes the attention weight, denotes a linear mapping of the node embedding vector; Step 225, embedding the enhanced node vector obtained in step 224 into the graph embedding vector obtained in step 223 The final node feature of each node is obtained by splicing, and is represented as: In Equation (12), denotes the final node feature, denotes the enhanced node embedding vector, denotes the graph embedding vector.
5. The deep reinforcement learning job shop scheduling method based on customer agent multi-bidding according to claim 1, wherein, The step three includes: Step 31, initialize the job shop agent; specifically, the job shop agent includes a feature extraction network and a calibration strategy network; Step 32, based on the customer agent of step two, obtain the feature representation of each process node in each customer agent using step 22; Step 33, input the feature representation of all process nodes obtained in step 32 into the calibration strategy network of the job shop agent, and the calibration strategy network is used as a decoder to output a calibration decision sequence in an autoregressive manner; Step 34, according to the calibration decision sequence obtained in step 33, execute g calibration schemes respectively, and calculate the reward value corresponding to the g calibration schemes; Step 35, after iterating steps 32-34 for B times, train the job shop agent using a policy gradient algorithm; specifically, the calculation expression of the policy gradient algorithm is: In Equation (13), denotes the gradient of the objective function; B denotes the number of iterations, denotes the reward value of the scheme , μ(τ b ) denotes the mean of the reward values of all schemes in the bth iteration, σ(τ b ) denotes the standard deviation of the reward values of all schemes in the bth iteration, ent coeff denotes the coefficient of the entropy, entropy denotes the entropy value, l denotes the lth calibration scheme, and m denotes the number of machines; Step 36, repeat steps 32-35 until the job shop agent policy converges.
6. The deep reinforcement learning job shop scheduling method based on customer agent multi-bidding according to claim 1, wherein, The step four includes: Step 41, load the customer agent trained in step two, and load the job shop agent trained in step three; specifically, load the parameters of the feature extraction network, the bidding strategy network and the state value network in the customer agent; load the parameters of the calibration strategy network in the job shop agent; Step 42, using the customer agent of step two, extract the feature representation of each process node for each customer agent respectively; Step 43, each client agent respectively bids the bidding strategy network of the client agent intelligent agent, taking the final node features of each process in the process set to be processed at the current decision moment as the input of the bidding strategy network, and the bidding strategy network outputs a respective bidding process set, thereby obtaining the bidding process set bid of all client agent intelligent agents set ; specifically, the bidding process set of all client agent intelligent agents is obtained by aggregating the bidding process set of each of the N client agent intelligent agents set , which is represented as: Step 44, the job shop agent intelligence body receives the bidding process set bid from all customer agent intelligence bodies set and the final node characteristics corresponding to each process, and performs autoregressive bidding decision by using the bidding strategy network to obtain a bidding decision result; wherein, during the execution of the autoregressive bidding decision, when the length of the bidding sequence corresponding to the autoregressive bidding decision reaches the number m of machines or all bidding processes have completed scheduling, the output of the autoregressive bidding decision performed by the bidding strategy network is stopped. Step 45, execute the calibration decision result obtained in step 44, and update the multi-agent job shop scheduling environment based on the job shop agent; Step 46, according to the target type of each client agent, the reward value of each client agent corresponding to the target type is calculated, and the calculation formula is: In Equation (14), represents a reward value, represents a number of scheduled processes of the customer agent CA i of the i-th customer, represents a value of a scheduled process node ; wherein the state value network of the customer agent outputs the value of the scheduled process node. Step 47, based on the value of the process node, calculate the reward value of the multi-agent job shop, and the calculation formula is: In Equation (15), represents the reward value of the multi-agent job shop, and N represents the number of customer agents, represents the number of scheduled processes of the customer agent CA i of the i-th customer, represents the value of the scheduled process node ; Step 48, repeat steps 42-47 until all processes of each customer agent are completed; Step 49, train the customer agent and the job shop agent using a multi-agent deep deterministic policy gradient algorithm; Step 410, repeat steps 42-49 until the customer agent and the job shop agent policy converges.
Citation Information
Patent Citations
Parallel machine multi-agent auction negotiation scheduling method
CN112101712A
Micro-service-multi-agent factory scheduling model based on deep reinforcement learning network
CN115437321A