Deep reinforcement learning job shop scheduling method based on customer agent multi-bidding
By adopting a deep reinforcement learning scheduling method based on customer agent multi-tender in the operation workshop, the problems of personalized goal satisfaction and agency preference protection in multi-agent operation workshop scheduling are solved, and efficient and personalized scheduling scheme generation is achieved.
Patent Information
- Application Number
- CN202510221851.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-27
AI Technical Summary
The prior art is difficult to effectively schedule multi-agent operation workshops, especially when multiple customer agent personalization goals need to be met, and the centralized optimization method cannot be applied to situations where agents are unwilling to disclose target preferences.
The deep reinforcement learning operation workshop scheduling method based on customer agent multi-tendering is adopted. Through the bidding mechanism, the customer agent agent bids multiple processes to be processed in each round. The operation workshop agent agent agent performs calibration decisions to ensure that the private preferences of each customer agent are protected and participate in the scheduling decisions.
Multiple agents with personalized goals have been achieved to participate in the scheduling decision-making, generate scheduling solutions that are acceptable to all parties, meet the private preferences of customer agents, and improve the effective utilization efficiency of operation workshop resources.
Smart Images

Figure CN120070018A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of job shop scheduling, and particularly relates to a deep reinforcement learning job shop scheduling method based on multi-bidding of customer agents. Background Art
[0002] Job shop scheduling is the core of discrete manufacturing enterprises to implement production management and ensure the orderly operation of the workshop. With the rapid development of the market economy, the enterprise production has changed from large-scale mass production to personalized customized production mode driven by user needs, resulting in the job shop scheduling changing from the original optimization goal mainly considering the optimization of job shop resources to a personalized goal considering and meeting the differentiated needs of users on the basis of workshop resource optimization. This has transformed the traditional job shop scheduling problem into a multi-agent job shop scheduling problem. The multi-agent job shop scheduling problem refers to modeling customers as customer agents and the job shop as a job shop agent, that is, multiple agents cooperate to complete the job shop scheduling problem. The main research content of the job shop scheduling problem is how to effectively allocate the workpieces from multiple competing customer agents on the limited machine resources owned by the job shop agent to form a scheduling plan acceptable to all parties.
[0003] The literature "Dong Z, Ren T, Qi F, et al. A reinforcement learning-based approach for solving multi-agent job shop scheduling problem. International Journal of Production Research, 2024: 1-26." discloses a multi-agent job shop scheduling optimization method based on deep reinforcement learning. This method combines a graph neural network and a Transformer model to propose a graph Transformer network structure, which can quickly generate high-quality scheduling plans. In the process of generating a scheduling plan by the above graph Transformer network structure, the target information of all customers needs to become the scheduling goal, and the job shop scheduling problem with multiple competing customer agents is solved in a centralized optimization manner according to this scheduling goal. All decisions in the scheduling optimization method are made by a central decision maker, focusing on optimizing the overall goal of the system. The scheduling plan obtained by the above graph Transformer network structure is difficult to meet the personalized goal requirements of multiple customer agents.
[0004] In addition, for reasons such as commercial competition or confidentiality, agents usually do not want to disclose their respective target preferences, resulting in the scheduling objectives based on all customer target information becoming private information of the agents. As a result, the method of centralized optimization for job shop scheduling is no longer applicable to such problems. Summary of the Invention
[0005] The object of the present invention is to solve the problem in the prior art that it is impossible to schedule a multi-agent job shop, and to provide a deep reinforcement learning job shop scheduling method based on multi-bidding of customer agents. As an effective resource allocation framework, in each round of bidding, the customer agent intelligent agents obtained by deep reinforcement learning respectively bid on multiple processes to be processed, and the job shop agent intelligent agent makes a decision on awarding the bid. And on the premise of being able to protect the private preferences of each customer agent, multiple agents with personalized goals are enabled to jointly participate in the scheduling decision-making to form a scheduling plan acceptable to all parties.
[0006] To achieve the above object, the technical solution provided by the present invention is as follows:
[0007] A deep reinforcement learning job shop scheduling method based on multi-bidding of customer agents, comprising:
[0008] Step 1: Model the multi-agent job shop scheduling problem based on customer agents and job shop agents, and obtain the customer agent intelligent agents corresponding to the customer agents, the job shop agent intelligent agents corresponding to the job shop, and obtain the scheduling state of the job shop according to the multi-agent job shop scheduling problem;
[0009] Step 2: Based on the scheduling state, use a feature extraction network to obtain the feature representations of each process node, use a bidding strategy network to output the set of processes participating in the bidding in each round of bidding, and use a proximal policy optimization algorithm to train the customer agent intelligent agents;
[0010] Step 3: Based on the feature representations of each process node obtained by the feature extraction network, use a bid-awarding strategy network to output the process scheduling order in each round of bidding as the bid-awarding decision order, and use a policy gradient algorithm to train the job shop agent intelligent agents;
[0011] Step 4: Jointly train the customer agent intelligent agents after training in Step 2 and the job shop agent intelligent agents after training in Step 3 using a multi-agent deep deterministic policy gradient algorithm; wherein, in the joint training algorithm, each customer agent intelligent agent only observes its own process information for multi-bidding, and the job shop agent intelligent agent collects the set of bidding processes of all customer agents to make a bid-awarding decision, and continuously repeats the above bidding process to generate a final scheduling plan.
[0012] As a further limitation of the present invention, the first step includes:
[0013] Step 11: Establish communication between the customer agent and the job shop agent, and determine the set of agents participating in the scheduling decision; specifically, the set of agents participating in the scheduling decision is expressed as:
[0014] {JSA, CA 1 , CA 2 , …, CA i , …, CA N} Formula (1)
[0015] In formula (1), JSA represents the job shop agent, and CA 1 represents the customer agent of the first customer, and CA 2 represents the customer agent of the second customer, and CA i represents the customer agent of the i-th customer, and CA N represents the customer agent of the N-th customer, and N represents the number of customer agents;
[0016] Step 12: Determine the set of machines of the job shop agent; specifically, the set of machines of the job shop agent is expressed as:
[0017] M = {M 1 , M 2 , …, M m} Formula (2)
[0018] In formula (2), M represents the set of machines of the job shop agent JSA, and M 1 represents the first machine, and M 2 represents the second machine, and M m represents the m-th machine, and m represents the number of machines;
[0019] Step 13: Obtain the set of workpieces to be processed by each customer agent; specifically, the set of workpieces to be processed by the customer agent CA of the i-th customer i is expressed as:
[0020]
[0021] In formula (3), represents the set of workpieces to be processed by the customer agent CA of the i-th customer i , represents the first workpiece to be processed in the customer agent CA of the i-th customer i , represents the second workpiece to be processed in the customer agent CA of the i-th customer i , represents the customer agent CA of the i-th customer iThe j-th workpiece to be processed in represents the customer agent CA of the i-th customer i the n-th i workpiece to be processed; where:
[0022] (1) For the j-th workpiece to be processed in the customer agent CA of the i-th customer i the completion time is and the due date is
[0023]
[0024] (2) In the set of workpieces to be processed of the customer agent CA of the i-th customer i the process information of each workpiece to be processed is expressed as:
[0025]
[0026] In formula (4), represents the customer agent CA of the i-th customer i the j-th workpiece to be processed in the k-th operation is processed on the machine and its processing time is where, j ∈ [1, n i , k ∈ [1, m];
[0027] Step 14. Determine the goal of the customer agent; specifically, the goal f of the customer agent CA of the i-th customer i includes: minimizing the makespan C i , minimizing the total completion time TC i and minimizing the total tardiness TT i , which is expressed as: i
[0028]
[0029] In formula (5), represents the completion time, n i represents a total of n i workpieces to be processed, represents the due date;
[0030] Step 15. Determine the goal of the job shop agent to minimize the makespan; specifically, minimizing the makespan C max The calculation expression is:
[0031]
[0032] In formula (6), i represents the i-th customer, and C i represents minimizing the makespan;
[0033] Step 16: Model the customer agent and the job shop agent respectively to obtain the customer agent intelligent body and the job shop agent intelligent body; specifically, the customer agent intelligent body is denoted as CA i , and the job shop agent intelligent body is denoted as JSA; among them, the customer agent intelligent body CA i makes a bidding decision according to the bidding strategy network, and the job shop agent intelligent body JSA executes a winning bid decision according to the winning bid strategy network;
[0034] Step 17: Construct a job shop scheduling environment and obtain the scheduling status of the job shop; specifically, according to the job shop scheduling environment, a disjunctive graph model is used to describe the scheduling status of the job shop; among them, the disjunctive graph model is expressed as:
[0035]
[0036] In formula (7), represents the disjunctive graph model, represents the set of operation nodes, represents the set of directed arcs of adjacent operations on the same workpiece, represents the set of undirected arcs of adjacent operations on the same machine.
[0037] As a further limitation of the present invention, the second step includes:
[0038] Step 21: Initialize the customer agent intelligent body; specifically, the customer agent intelligent body includes a feature extraction network, a bidding strategy network, and a state value network; the feature extraction network includes a graph neural network and an encoder network; both the bidding strategy network and the state value network adopt a fully connected neural network structure;
[0039] Step 22: The feature extraction network in the customer agent intelligent body obtains the feature representation of each operation node; specifically, first, the graph neural network in the feature extraction network is used to extract the feature representation of each operation node in the disjunctive graph model; then, the feature representation of each operation node is input into the encoder network in the feature extraction network, and the final feature representation of each operation node is obtained based on the encoder network;
[0040] Step 23: Input the feature representations of all the operation nodes obtained in Step 22 into the bidding strategy network, and output the probability of selecting each operation node based on the bidding strategy network; specifically, m currently processable operations are selected and entered into the bidding operation set in descending order of the probability value, denoted as If the number of currently processable operations is less than m, then all processable operations are added to the tender operation set
[0041] Step 24: Input the feature representations of all processable operation nodes obtained in Step 22 into the state value network, and obtain the estimated values of all currently processable operations based on the state value network
[0042] Step 25: Execute the tender operation set obtained in Step 23 Update the job shop scheduling environment state s t ;
[0043] Step 26: Obtain the target type based on the goal of the customer agent, and calculate the reward value through the target type Specifically, the reward value is calculated by the following formula
[0044]
[0045] In formula (8), represents the customer agent CA corresponding to the i-th customer i The reward value of f i (s t ) represents the objective function value of the customer agent CA in state s t ; f i The objective function value of the customer agent CA in state s i (s t+1 ) represents the objective function value of the customer agent CA in state s t+1 ; i The objective function value of
[0046] Step 27: Repeat Steps 22 - 26 until all operations are scheduled;
[0047] Step 28: Train the customer agent intelligent agent using the proximal policy optimization algorithm;
[0048] Step 29: Repeat Steps 22 - 28 until the customer agent intelligent agent policy converges.
[0049] As a further limitation of the present invention, Step 22 includes:
[0050] Step 221: The features of the operation nodes in the disjunctive graph model include: the estimated earliest completion time of the operation nodes Whether the scheduling is completed
[0051] Step 222: The graph neural network performs K iterations to extract the embedding vectors of the operation nodes The calculation formula for the q-th round of iteration is as follows:
[0052]
[0053] In formula (9), represents the embedding vector of process node in the q-th round of iteration, and MLP (q) represents the fully connected neural network of the q-th layer of the graph neural network. ∈ represents a learnable parameter. represents the original feature of process node ; represents the set of neighborhood nodes of process node ; represents the embedding vector of neighborhood node u in the (q - 1)-th round of iteration;
[0054] Step 223: Extract the embedding vector from the graph neural network and perform average pooling on it to obtain the graph embedding vector Among them, the graph embedding vector The calculation formula is as follows:
[0055]
[0056] In formula (10), n i represents a total of n i workpieces to be processed, and m represents the number of machines. represents the process node, represents the set of process nodes, represents the process node 's node embedding vector;
[0057] Step 224: Input the node embedding vector obtained in Step 222 into the encoder network, and obtain the enhanced node embedding vector based on the encoder network. The expression is:
[0058]
[0059] In formula (11), represents the set of neighborhood nodes of process node ; represents the finally obtained enhanced node embedding vector, represents the attention weight, represents the linear mapping of the node embedding vector;
[0060] Step 225: Concatenate the enhanced node embedding vector obtained in Step 224 and the graph embedding vector obtained in Step 223 to obtain the final node feature of each node, which is expressed as:
[0061]
[0062] In formula (12), represents the final node feature, represents the enhanced node embedding vector, represents the graph embedding vector.
[0063] As a further limitation of the present invention, step three includes:
[0064] Step 31, initialize the job shop agent intelligent agent; specifically, the job shop agent intelligent agent includes a feature extraction network and a calibration policy network;
[0065] Step 32, based on the customer agent intelligent agent in step two, use step 22 to obtain the feature representations of each process node in each customer agent;
[0066] Step 33, input the feature representations of all process nodes obtained in step 32 into the calibration policy network of the job shop agent intelligent agent, and the calibration policy network outputs a calibration decision sequence in an autoregressive manner as a decoder;
[0067] Step 34, according to the calibration decision sequence obtained in step 33, execute m calibration schemes respectively, and calculate the reward values corresponding to the m calibration schemes;
[0068] Step 35, after iteratively executing step 32 - step 34 for B times, train the job shop agent intelligent agent using the policy gradient algorithm; specifically, the calculation expression of the policy gradient algorithm is:
[0069]
[0070] In formula (13), represents the gradient of the objective function; B represents the number of iterations, m represents the number of machines, represents the scheme 's reward value, μ(τ b ) represents the mean of the reward values of all schemes in the b-th iteration, σ(τ b ) represents the standard deviation of the reward values of all schemes in the b-th iteration, ent coeff represents the coefficient of entropy, and entropy represents the entropy value;
[0071] Step 36, repeatedly execute step 32 - step 35 until the policy of the job shop agent intelligent agent converges.
[0072] As a further limitation of the present invention, step four includes:
[0073] Step 41: Load the customer agent intelligent agent trained in Step 2 and load the job shop agent intelligent agent trained in Step 3; specifically, load the parameters of the feature extraction network, bidding strategy network, and state value network in the customer agent intelligent agent; load the parameters of the calibration strategy network in the job shop agent intelligent agent.
[0074] Step 42: Use the customer agent intelligent agent in Step 2 to extract the feature representations of each process node for each customer agent's respective customer agent intelligent agent.
[0075] Step 43: For the bidding strategy network of each customer agent's respective customer agent intelligent agent, use the final node features of each process in the set of processes to be processed at the current decision-making moment as the input of the bidding strategy network. The bidding strategy network outputs its respective set of bidding processes, and the set of bidding processes bid of all customer agent intelligent agents is obtained; specifically, summarize the sets of bidding processes of each of the N customer agent intelligent agents to obtain the set of bidding processes bid of all customer agent intelligent agents. set ; specifically, summarize the sets of bidding processes of each of the N customer agent intelligent agents to obtain the set of bidding processes bid of all customer agent intelligent agents. set which is expressed as:
[0076] Step 44: The job shop agent intelligent agent receives the set of bidding processes bid from all customer agent intelligent agents set , and the corresponding final node features of each process, and uses the calibration strategy network to execute an autoregressive calibration decision to obtain a calibration decision result; where, during the execution of the autoregressive calibration decision, when the length of the calibration sequence corresponding to the autoregressive calibration decision reaches the number m of machines or all the bidding processes have been scheduled, stop the output of the autoregressive calibration decision of the calibration strategy network.
[0077] Step 45: Execute the calibration decision result obtained in Step 44, and update the multi-agent job shop scheduling environment based on the job shop agent intelligent agent.
[0078] Step 46: Calculate the reward value corresponding to the target type of each customer agent according to the target type of each customer agent. The calculation formula is:
[0079]
[0080] In formula (14), represents the reward value, represents the customer agent CA of the i-th customer i the number of scheduled processes, represents the scheduled process node The value; among them, the state value network of the customer agent intelligent body outputs the value of the scheduled process node;
[0081] Step 47, calculate the reward value of the multi-agent job shop based on the value of the process node, and its calculation formula is:
[0082]
[0083] In formula (15), represents the reward value of the multi-agent job shop, N represents the number of customer agents, represents the customer agent CA of the i-th customer i the number of scheduled processes, represents the scheduled process node the value;
[0084] Step 48, repeatedly execute Step 42 - Step 47 until all processes of each customer agent are scheduled.
[0085] Step 49, use the multi-agent deep deterministic policy gradient algorithm to train the customer agent intelligent body and the job shop agent intelligent body;
[0086] Step 410, repeatedly execute Step 42 - Step 49 until the strategies of the customer agent intelligent body and the job shop agent intelligent body converge.
[0087] The advantages of the present invention are:
[0088] 1. The multi-agent job shop scheduling method of the present invention schedules all processes of the customer agent based on deep reinforcement learning and the bidding mechanism. Compared with the traditional scheduling method, the present invention not only allows the customer agent and the resource agent of the job shop to jointly participate in the scheduling decision-making. In each round of bidding, the customer agent intelligent body trained by deep reinforcement learning bids for multiple processes to be processed respectively, and the job shop agent intelligent body makes the decision of awarding the bid, realizing multiple agents with personalized goals to jointly participate in the scheduling decision-making.
[0089] 2. The present invention takes the personalized needs of the customer agent as the core consideration factor in the scheduling decision-making process, so that the scheduling scheme generated by the present invention fully reflects the preference information of each customer agent; the present invention adopts the method that each customer agent bids multiple times in each round of bidding, greatly meeting the personalized needs of the customer agent to fully express its private preferences.
[0090] 3. The job shop scheduling method of the present invention combines deep reinforcement learning with a bidding mechanism. The bidding mechanism ensures the fair allocation of job shop resources among customer agents, and deep reinforcement learning enables customer agents and job shop agents to autonomously learn resource conflict negotiation methods in multiple rounds of bidding, ensuring the effective utilization of job shop resources. In addition, the customer agent intelligent body and the job shop agent intelligent body of the present invention are trained by a joint training algorithm, and have good computational efficiency and solution performance.
[0091] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0092] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the description of the embodiments in conjunction with the following drawings, in which:
[0093] Figure 1 : Flowchart of a deep reinforcement learning job shop scheduling method based on multi-bidding of customer agents provided by the present invention;
[0094] Figure 2 : Schematic diagram of the bidding algorithm of the customer agent intelligent body provided by the present invention;
[0095] Figure 3 : Schematic diagram of the tender awarding algorithm of the job shop agent intelligent body provided by the present invention;
[0096] Figure 4 : Schematic diagram of the joint training algorithm of the customer agent intelligent body and the job shop agent intelligent body provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0097] The embodiments of the present invention will be described in detail below. The embodiments are exemplary and are intended to explain the present invention, but should not be construed as limiting the present invention.
[0098] Please refer to Figure 1 , the embodiment of the present invention provides a deep reinforcement learning job shop scheduling method based on multi-bidding of customer agents, including:
[0099] Step 1: Model the multi-agent job shop scheduling problem based on customer agents and job shop agents, and obtain the customer agent intelligent body corresponding to the customer agent, the job shop agent intelligent body corresponding to the job shop, and the scheduling state of the job shop according to the multi-agent job shop scheduling problem.
[0100] Step 1 of the embodiment of the present invention includes:
[0101] Step 11: Establish communication between the customer agent and the job shop agent to determine the set of agents participating in the scheduling decision. Specifically, the set of agents participating in the scheduling decision is represented as:
[0102] {JSA, CA 1 , CA 2 , …, CA i , …, CA N} Formula (1)
[0103] In Formula (1), JSA represents the job shop agent, and CA 1 represents the customer agent of the first customer, CA 2 represents the customer agent of the second customer, CA i represents the customer agent of the i-th customer, CA N represents the customer agent of the N-th customer, and N represents the number of customer agents;
[0104] Step 12: Determine the set of machines of the job shop agent. Specifically, the set of machines of the job shop agent is represented as:
[0105] M = {M 1 , M 2 , …, M m} Formula (2)
[0106] In Formula (2), M represents the set of machines of the job shop agent JSA, and M 1 represents the first machine, M 2 represents the second machine, M m represents the m-th machine, and m represents the number of machines;
[0107] Step 13: Obtain the set of workpieces to be processed by each customer agent. Specifically, the set of workpieces to be processed by the customer agent CA i of the i-th customer is represented as:
[0108]
[0109] In Formula (3), represents the set of workpieces to be processed by the customer agent CA i of the i-th customer, represents the first workpiece to be processed in the customer agent CA i of the i-th customer, represents the second workpiece to be processed in the customer agent CA i of the i-th customer, represents the j-th workpiece to be processed in the customer agent CA i of the i-th customer, represents the customer agent CA ithe nth in i workpieces to be processed; where:
[0110] (1) For the jth workpiece to be processed in the customer agent CA of the ith customer i the completion time is and the due date is
[0111]
[0112] (2) In the set of workpieces to be processed of the customer agent CA of the ith customer i the process information of each workpiece to be processed is expressed as:
[0113]
[0114] In formula (4), represents the jth workpiece to be processed in the customer agent CA of the ith customer i the kth process is processed on the machine and its processing time is where j ∈ [1, n i , k ∈ [1, m];
[0115] Step 14. Determine the goal of the customer agent; specifically, the goal f i of the customer agent CA of the ith customer i includes: minimizing the makespan C i , minimizing the total completion time TC i and minimizing the total tardiness TT i , expressed as:
[0116]
[0117] In formula (5), represents the completion time, n i represents a total of n i workpieces to be processed, represents the due date;
[0118] Step 15. Determine the goal of the job shop agent to minimize the makespan; specifically, the calculation expression for minimizing the makespan C max is:
[0119]
[0120] In formula (6), i represents the ith customer, and C i represents minimizing the makespan;
[0121] Step 16: Model the customer agent and the job shop agent respectively to obtain the customer agent intelligent agent and the job shop agent intelligent agent. Specifically, the customer agent intelligent agent is denoted as CA i , and the job shop agent intelligent agent is denoted as JSA. Among them, the customer agent intelligent agent CA i makes a bidding decision according to the bidding strategy network, and the job shop agent intelligent agent JSA executes a winning bid decision according to the winning bid strategy network;
[0122] Step 17: Construct a job shop scheduling environment and obtain the scheduling status of the job shop. Specifically, according to the job shop scheduling environment, a disjunctive graph model is used to describe the scheduling status of the job shop. Among them, the disjunctive graph model is expressed as:
[0123]
[0124] In formula (7), represents the disjunctive graph model, represents the set of process nodes, represents the set of directed arcs of adjacent processes on the same workpiece, represents the set of undirected arcs of adjacent processes on the same machine.
[0125] Step 2: Based on the scheduling status, use a feature extraction network to obtain the feature representation of each process node, use a bidding strategy network to output the set of processes participating in the bidding in each round of bidding, and use the proximal policy optimization algorithm to train the customer agent intelligent agent.
[0126] Step 2 of the embodiment of the present invention includes:
[0127] Step 21: Initialize the customer agent intelligent agent. Specifically, the customer agent intelligent agent includes a feature extraction network, a bidding strategy network, and a state value network. The feature extraction network includes a graph neural network and an encoder network. Both the bidding strategy network and the state value network adopt a fully connected neural network structure;
[0128] Step 22: The feature extraction network in the customer agent intelligent agent obtains the feature representation of each process node. Specifically, first, use the graph neural network in the feature extraction network to extract the feature representation of each process node in the disjunctive graph model; then, input the feature representation of each process node into the encoder network in the feature extraction network, and based on the encoder network, obtain the final feature representation of each process node;
[0129] Step 23: Input the feature representations of all process nodes obtained in Step 22 into the bidding strategy network, and based on the bidding strategy network, output the probability of selecting each process node. Specifically, in the order from large to small according to the probability value, select m currently processable processes to enter the bidding process set, denoted as If the number of currently processable operations is less than m, all processable operations are added to the tender operation set
[0130] Step 24: Input the feature representations of all processable operation nodes obtained in Step 22 into the state-value network, and obtain the estimated values of all currently processable operations based on the state-value network
[0131] Step 25: Execute the tender operation set obtained in Step 23 Update the job shop scheduling environment state s t ;
[0132] Step 26: Obtain the target type based on the target of the customer agent, and calculate the reward value through the target type Specifically, the reward value is calculated by the following formula
[0133]
[0134] In formula (8), represents the customer agent CA corresponding to the i-th customer i The reward value of f i (s t ) represents the objective function value of the customer agent CA in state s t ; f i (s i (s t+1 ) represents the objective function value of the customer agent CA in state s t+1 ; i The objective function value of the customer agent CA;
[0135] Step 27: Repeat Steps 22 - 26 until all operations are scheduled;
[0136] Step 28: Train the customer agent intelligent body using the proximal policy optimization algorithm
[0137] Step 29: Repeat Steps 22 - 28 until the policy of the customer agent intelligent body converges
[0138] More specifically, the above Step 22 of the embodiment of the present invention includes:
[0139] The features of the operation nodes in the disjunctive graph model include: the estimated earliest completion time of the operation nodes Whether the scheduling is completed
[0140] Step 221: The features of the operation nodes in the disjunctive graph model include: the estimated earliest completion time of the operation nodes The calculation formula for the q-th iteration is:
[0141]
[0142] In formula (9), represents the process node 's embedding vector in the q-th round of iteration, and MLP (q) represents the fully connected neural network of the q-th layer of the graph neural network, and ∈ represents a learnable parameter. represents the process node 's original feature. represents the process node 's set of neighborhood nodes. represents the embedding vector of the neighborhood node u in the (q - 1)-th round of iteration;
[0143] Step 223: Extract the embedding vector from the graph neural network and perform average pooling on it to obtain the graph embedding vector. Among them, the graph embedding vector The calculation formula is:
[0144]
[0145] In formula (10), n i represents a total of n i workpieces to be processed, m represents the number of machines. represents the process node, represents the set of process nodes, represents the process node 's node embedding vector;
[0146] Step 224: Input the node embedding vector obtained in Step 222 into the encoder network, and obtain the enhanced node embedding vector based on the encoder network. The expression is:
[0147]
[0148] In formula (11), represents the set of neighborhood nodes of the process node , represents the finally obtained enhanced node embedding vector, represents the attention weight, represents a linear mapping of the node embedding vector;
[0149] Step 225: Concatenate the enhanced node embedding vector obtained in Step 224 and the graph embedding vector obtained in Step 223 to obtain the final node feature of each node, which is expressed as:
[0150]
[0151] In formula (12), represents the final node feature, represents the enhanced node embedding vector, represents the graph embedding vector.
[0152] Please refer to Figure 2 For the above customer agent intelligent agent in the embodiment of the present invention, in practical applications, first, initialize the job shop scheduling environment and the feature extraction network, bidding strategy network, and state value network of the customer agent intelligent agent. The feature extraction network includes a graph neural network and an encoder network. The feature extraction network is used to obtain the feature representations of each process node. Specifically, the graph neural network is adopted in the feature extraction network to obtain the embedding vectors of each process node, and average pooling is performed on the node embedding vectors to obtain the graph embedding vector. The node embedding vectors are input into the encoder network to obtain the encoder embedding vectors, and the graph embedding vector and the encoder embedding vector are concatenated to obtain the final node feature. The final node feature of the currently processable process is input into the bidding strategy network to output the probabilities of selecting each process, and at most m currently processable processes are selected in descending order of the probability values to enter the bidding process set; the final node feature of the currently processable process is input into the state value network to obtain its state value. Execute the bidding process set and update the job shop scheduling environment. Repeat the above operations until all processes are scheduled. After all processes of the customer agent are scheduled, the proximal policy optimization algorithm is used to train the customer agent intelligent agent.
[0153] Step 3: Based on the feature representations of each process node obtained by the feature extraction network, use the calibration strategy network to output the process scheduling order in each round of bidding as the calibration decision order, and use the policy gradient algorithm to train the job shop agent intelligent agent.
[0154] Specifically, the above Step 3 in the embodiment of the present invention includes:
[0155] Step 31: Initialize the job shop agent intelligent agent; specifically, the job shop agent intelligent agent includes a feature extraction network and a calibration strategy network;
[0156] Step 32: Based on the customer agent intelligent agent in Step 2, use Step 22 to obtain the feature representations of each process node in each customer agent;
[0157] Step 33: Input the feature representations of all process nodes obtained in Step 32 into the calibration strategy network of the job shop agent intelligent agent, and the calibration strategy network outputs the calibration decision sequence in an autoregressive manner as a decoder;
[0158] Step 34: According to the calibration decision sequence obtained in Step 33, execute m calibration schemes respectively, and calculate the reward values corresponding to the m calibration schemes;
[0159] Step 35: After iteratively executing Steps 32 - 34 for B times, train the job - shop agent intelligent agent using the policy gradient algorithm; specifically, the calculation expression of the policy gradient algorithm is:
[0160]
[0161] In formula (13), represents the gradient of the objective function; B represents the number of iterations, m represents the number of machines, represents the solution 's reward value, μ(τ b ) represents the mean of the reward values of all solutions in the b - th iteration, σ(τ b ) represents the standard deviation of the reward values of all solutions in the b - th iteration, ent coeff represents the coefficient of entropy, and entropy represents the entropy value;
[0162] Step 36: Repeat Steps 32 - 35 until the policy of the job - shop agent intelligent agent converges.
[0163] Initialize the job - shop scheduling environment, the feature extraction network, and the calibration policy network of the job - shop agent. Concatenate the encoder embedding vector and the graph embedding vector obtained by the feature extraction network as the final node feature. The final node feature of each node is used as the key vector and the value vector to input into the multi - pointer network, and the final node feature of each node and the multi - scheduling sequence together constitute the enhanced state embedding vector, which is used as the query vector to input into the multi - pointer network. The multi - pointer network receives the query vector, the key vector, and the value vector, and its output result passes through a linear layer and a Softmax layer in sequence to obtain the probability of selecting each operation to be processed. According to this probability, generate m scheduling sequences to balance the mean and variance in the training process. Execute the m scheduling sequences respectively, and update the enhanced state embedding vector. Continuously repeat this process until all operations of the job - shop agent are scheduled. After all operations are scheduled, train the job - shop agent intelligent agent using the policy gradient algorithm.
[0164] Step Four: For the customer agent intelligent agent after training in Step Two and the job - shop agent intelligent agent after training in Step Three, conduct joint training using the multi - agent deep deterministic policy gradient algorithm; among them, in the joint training algorithm, each customer agent intelligent agent only observes its own operation information to make multiple bids, and the job - shop agent intelligent agent collects the set of bid operations of all customer agents to make calibration decisions. Continuously repeat the above bidding process to generate the final scheduling plan.
[0165] Specifically, step four of the embodiments of the present invention includes:
[0166] Step 41: Load the customer agent intelligent agent trained in step two and the job shop agent intelligent agent trained in step three. Specifically, load the parameters of the feature extraction network, bidding strategy network, and state value network in the customer agent intelligent agent; load the parameters of the calibration strategy network in the job shop agent intelligent agent;
[0167] Step 42: Use the customer agent intelligent agent in step two to extract the feature representations of each process node for the respective customer agent intelligent agents of each customer agent.
[0168] Step 43: The bidding strategy network of the respective customer agent intelligent agents of each customer agent uses the final node features of each process in the set of processes to be processed at the current decision-making moment as the input of the bidding strategy network. The bidding strategy network outputs their respective sets of bidding processes to obtain the set of bidding processes bid of all customer agent intelligent agents. set Specifically, summarize the sets of bidding processes of the respective customer agent intelligent agents of N customer agents to obtain the set of bidding processes bid of all customer agent intelligent agents. set It is expressed as:
[0169] Step 44: The job shop agent intelligent agent receives the set of bidding processes bid from all customer agent intelligent agents set , and the corresponding final node features of each process, and uses the calibration strategy network to execute the autoregressive calibration decision to obtain the calibration decision result. Among them, during the execution of the autoregressive calibration decision, when the length of the calibration sequence corresponding to the autoregressive calibration decision reaches the number m of machines or all the bidding processes have been scheduled, stop the output of the autoregressive calibration decision of the calibration strategy network.
[0170] Step 45: Execute the calibration decision result obtained in step 44, and update the multi-agent job shop scheduling environment based on the job shop agent intelligent agent.
[0171] Step 46: Calculate the reward value corresponding to the target type of each customer agent according to the target type of each customer agent. The calculation formula is:
[0172]
[0173] In formula (14), represents the reward value, represents the number of scheduled processes of the customer agent CA of the i-th customer i , represents the value of the scheduled process node . Among them, the state value network of the customer agent intelligent agent outputs the value of the scheduled process node.
[0174] Step 47: Calculate the reward value of the multi-agent job shop based on the value of the process node. The calculation formula is as follows:
[0175]
[0176] In formula (15), represents the reward value of the multi-agent job shop, N represents the number of customer agents, represents the customer agent CA of the i-th customer i The number of scheduled processes,[ represents the scheduled process node Value;
[0177] Step 48: Repeat steps 42 - 47 until all processes of each customer agent are scheduled;
[0178] Step 49: Use the multi-agent deep deterministic policy gradient algorithm to train the customer agent and the job shop agent;
[0179] Step 410: Repeat steps 42 - 49 until the policies of the customer agent and the job shop agent converge.
[0180] Initialize the job shop scheduling environment, and load the customer agent trained in step 2 and the job shop agent trained in step 3. The feature extraction network of the customer agent obtains the final node features of each process. Among them, all customer agents share a graph neural network to improve the calculation efficiency. The bidding strategy network of each customer agent takes the final node features of its respective processes to be processed as input and outputs its respective set of bidding processes. The job shop agent receives the set of bidding processes and their corresponding final node features from all customer agents, and uses the calibration strategy network to perform autoregressive calibration decisions. During the calibration decision process, when the calibration sequence length reaches the number of machines or all bidding processes have been scheduled, stop the autoregressive output to obtain the scheduling sequence. Execute the scheduling sequence to update the job shop scheduling environment state. Continuously repeat the above process until all processes of all customer agents are scheduled. After all processes are scheduled, use the multi-agent deep deterministic policy gradient algorithm to train the customer agent and the job shop agent.
[0181] It should be noted that in the multi-bidding of the customer agent in the embodiments of the present invention, in each round of bidding, the customer agent can bid on multiple items, corresponding to the situation where only one item can be bid in each round. The agent in the embodiments of the present invention is a modeling technique that simulates and solves problems in the real world by abstracting entities in a complex system into individuals with autonomy, interactivity, and intelligence. Each agent has its own goals, behaviors, and decision-making capabilities, and can interact with other agents to jointly complete the scheduling task. The customer agent in the embodiments of the present invention is a term in a common scheduling problem, representing the description of the customer's needs and the activities of the interaction between the customer's needs and the job shop using agent technology.
[0182] A deep reinforcement learning job shop scheduling method based on multi-bidding of customer agents disclosed in the embodiments of the present invention. In each round of bidding, the customer agent intelligent agents obtained by deep reinforcement learning training bid on multiple processes to be processed respectively, and the job shop agent intelligent agent makes a tender award decision, enabling multiple agents with personalized goals to jointly participate in the scheduling decision. Specifically, the bidding strategy network of the customer agent intelligent agent in the embodiments of the present invention takes its private workpiece information as input at each decision moment and outputs the processes participating in the bidding; the tender award strategy network of the job shop agent intelligent agent takes the bidding process information as input, outputs the tender award decision result, and submits the tender award decision result to the job shop scheduling environment to update the disjunctive graph model. The above bidding process is repeated until the scheduling of the processes of all customer agents is completed. The job shop scheduling method in the embodiments of the present invention combines deep reinforcement learning with the bidding mechanism, can effectively solve the problem of collaborative scheduling of multiple agents with personalized goals, and has good computational efficiency and optimization effect. In addition, compared with the traditional scheduling method, the embodiments of the present invention allow the customer agent and the resource agent to jointly participate in the scheduling decision, and take its own personalized needs as the only consideration factor in the decision-making process, so that the generated scheduling plan can fully reflect the preference information of each agent; the combination of deep reinforcement learning and the bidding mechanism greatly improves the resource allocation efficiency and has good computational efficiency and solution performance; the customer agent in the embodiments of the present invention makes multiple bids in each round of bidding, fully expressing the private preferences of the customer agent.
[0183] As described above, the above are only the specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention.
Claims
1. A deep reinforcement learning job shop scheduling method based on customer agent multi-bidding, characterized in that: include: Step 1: Model a multi-agent job shop scheduling problem based on the customer agent and the job shop agent, and obtain the customer agent agent corresponding to the customer agent, the job shop agent agent corresponding to the job shop, and the scheduling status of the job shop according to the multi-agent job shop scheduling problem; Step 2: Based on the scheduling state, a feature extraction network is used to obtain feature representations of each process node, a bidding strategy network is used to output a set of processes participating in bidding in each round of bidding, and a proximal strategy optimization algorithm is used to train a customer agent; Step 3: Based on the feature representation of each process node obtained by the feature extraction network, a calibration strategy network is used to output the process scheduling sequence in each round of bidding as the calibration decision sequence, and a policy gradient algorithm is used to train the job shop agent; Step 4: The customer agent trained in step 2 and the job shop agent trained in step 3 are jointly trained using a multi-agent deep deterministic policy gradient algorithm; wherein, in the joint training algorithm, each customer agent only observes its own process information for multiple biddings, and the job shop agent collects the bidding process sets of all customer agents for calibration decisions, and the above bidding process is repeated continuously to generate a final scheduling plan.
2. According to claim 1, a deep reinforcement learning job shop scheduling method based on customer agent multi-bidding is characterized in that: The step one comprises: Step 11: Establish communication between the client agent and the job shop agent, and determine the set of agents participating in the scheduling decision. Specifically, the set of agents participating in the scheduling decision is expressed as: {JSA,CA1,CA2,…,CA i ,…,CA N }Formula (1) In formula (1), JSA represents the job shop agent, CA1 represents the customer agent of the first customer, CA2 represents the customer agent of the second customer, and CA i represents the customer agent of the ith customer, CA N represents the customer agent of the Nth customer, where N represents the number of customer agents; Step 12: Determine the machine set of the job shop agent; specifically, the machine set of the job shop agent is expressed as: M={M1,M2,…,M m } Formula (2) In formula (2), M represents the machine set of the job shop agent JSA, M1 represents the first machine, M2 represents the second machine, and M m represents the mth machine, where m represents the number of machines; Step 13: Obtain the set of workpieces to be processed of each customer agent; specifically, the customer agent CA of the i-th customer i The set of workpieces to be processed is expressed as: In formula (3), The client agent CA of the i-th client i A collection of workpieces to be processed, The client agent CA of the i-th client i The first workpiece to be processed, The client agent CA of the i-th client i The second workpiece to be processed, The client agent CA of the i-th client i The jth workpiece to be processed in The client agent CA of the i-th client i The nth i workpieces to be processed; among them: (1) The client agent CA of the i-th client i The jth workpiece to be processed The completion time is Delivery time is (2) The client agent CA of the i-th client i A collection of workpieces to be processed In the figure, the process information of each workpiece to be processed is expressed as: In formula (4), The client agent CA of the i-th client i The jth workpiece to be processed The kth process In the machine The processing time is Among them, j∈[1,n i ], k∈[1,m]; Step 14: Determine the target of the customer agent; specifically, the customer agent CA of the i-th customer i The goal of i Includes: Minimize the maximum completion time C i , minimize the total completion time TC i and minimize the total tardiness TT i , expressed as: In formula (5), represents the completion time, n i Indicates total n i Workpieces to be processed, Indicates the delivery period; Step 15: Determine the job shop agent's goal of minimizing the maximum completion time; specifically, minimize the maximum completion time C max The calculation expression is: In formula (6), i represents the i-th customer, C i It means minimizing the maximum completion time; Step 16: Model the customer agent and the job shop agent respectively to obtain the customer agent intelligent agent and the job shop agent intelligent agent; specifically, the customer agent intelligent agent is represented by CA i , the job shop agent is represented as JSA; among them, the customer agent CA i Bidding decisions are made according to the bidding strategy network, and the job shop agent JSA executes the bidding decision according to the bidding strategy network; Step 17: construct a job shop scheduling environment and obtain the scheduling status of the job shop; specifically, according to the job shop scheduling environment, a disjunctive graph model is used to describe the scheduling status of the job shop; wherein the disjunctive graph model is expressed as: In formula (7), represents the disjunctive graph model, Represents a set of process nodes, represents the set of directed arcs of adjacent processes on the same workpiece, A set of undirected arcs representing adjacent processes on the same machine.
3. The method for deep reinforcement learning job shop scheduling based on client agent multi-bidding according to claim 1, characterized in that: The second step comprises: Step 21, initializing the customer agent intelligent agent; specifically, the customer agent intelligent agent includes a feature extraction network, a bidding strategy network and a state value network; the feature extraction network includes a graph neural network and an encoder network; the bidding strategy network and the state value network both adopt a fully connected neural network structure; Step 22, the feature extraction network in the client agent agent obtains the feature representation of each process node; specifically, first, the graph neural network in the feature extraction network is used to extract the feature representation of each process node in the disjunctive graph model; then, the feature representation of each process node is input into the encoder network in the feature extraction network, and the final feature representation of each process node is obtained based on the encoder network; Step 23: Input the feature representations of all process nodes obtained in step 22 into the bidding strategy network, and output the probability of selecting each process node based on the bidding strategy network; specifically, select m currently processable processes into the bidding process set in descending order of probability values, expressed as If the number of currently processable processes is less than m, all processable processes will be added to the bidding process set. Step 24: Input the feature representations of all process nodes obtained in step 22 into the state value network, and obtain the estimated values of all current process nodes based on the state value network. Step 25: Execute the bidding process set obtained in step 23 Update the job shop scheduling environment status t ; Step 26: Get a target type based on the target of the client agent, and calculate the reward value based on the target type Specifically, the reward value is calculated by the following formula: In formula (8), Represents the customer agent CA corresponding to the i-th customer i Reward value f i (s t ) indicates state s t Download the client agent CA i The objective function value of i (s t+1 ) indicates state s t+1 Download the client agent CA i The objective function value of Step 27: Repeat steps 22 to 26 until all processes are scheduled; Step 28: Use the proximal strategy optimization algorithm to train the customer agent; Step 29: Repeat steps 22 to 28 until the client agent strategy converges.
4. The method for deep reinforcement learning job shop scheduling based on client agent multi-bidding according to claim 3 is characterized in that: The step 22 comprises: Step 221: The characteristics of the process nodes in the disjunctive graph model include: the estimated earliest completion time of the process node Is the scheduling completed? Step 222: The graph neural network performs K iterations to extract the embedding vector of the process node. The calculation formula for the qth iteration is: In formula (9), Indicates process node The embedding vector at the qth iteration, MLP (q) represents the qth layer of the graph neural network, ∈ represents a learnable parameter, Indicates process node The original characteristics of Indicates process node The set of neighboring nodes of Represents the embedding vector of the neighborhood node u in the q-1th round of iteration; Step 223: extract the embedding vector from the graph neural network and perform average pooling on it to obtain the graph embedding vector Among them, the graph embedding vector The calculation formula is: In formula (10), n i Indicates total n i workpieces to be processed, m represents the number of machines, Indicates the process node, Represents a set of process nodes, Indicates process node The node embedding vector of Step 224: embed the node obtained in step 222 into a vector The encoder network is input, and the enhanced node embedding vector is obtained based on the encoder network, and the expression is: In formula (11), Indicates process node The set of neighboring nodes of represents the final enhanced node embedding vector, represents the attention weight, Represents a linear mapping of node embedding vectors; Step 225: The enhanced node embedding vector obtained in step 224 and the graph embedding vector obtained in step 223 are added together. The final node features of each node are obtained by splicing, which is expressed as: In formula (12), represents the final node feature, represents the enhanced node embedding vector, Represents the graph embedding vector.
5. The method for deep reinforcement learning job shop scheduling based on client agent multi-bidding according to claim 1, characterized in that: The step three comprises: Step 31, initializing the job shop agent agent; specifically, the job shop agent agent includes a feature extraction network and a calibration strategy network; Step 32: Based on the client agent agent in step 2, obtain the feature representation of each process node in each client agent using step 22; Step 33: input the feature representations of all process nodes obtained in step 32 into the calibration strategy network of the job shop agent, and the calibration strategy network serves as a decoder to output a calibration decision sequence in an autoregressive manner; Step 34: According to the calibration decision sequence obtained in step 33, m calibration schemes are respectively executed, and the reward values corresponding to the m calibration schemes are calculated; Step 35, after iteratively executing step 32-step 34 for B times, use the policy gradient algorithm to train the job shop agent; specifically, the calculation expression of the policy gradient algorithm is: In formula (13), represents the gradient of the objective function; B represents the number of iterations, m represents the number of machines, Representation scheme The reward value, μ(τ b ) represents the mean reward value of all solutions in the bth iteration, σ(τ b ) represents the standard deviation of the reward values of all solutions in the b-th iteration, ent coeff It represents the coefficient of entropy, and entropy represents the entropy value; Step 36: Repeat steps 32 to 35 until the job shop agent strategy converges.
6. The method for deep reinforcement learning job shop scheduling based on client agent multi-bidding according to claim 1, characterized in that: The fourth step comprises: Step 41, loading the customer agent trained in step 2, loading the job shop agent trained in step 3; specifically, loading the parameters of the feature extraction network, bidding strategy network and state value network in the customer agent; loading the parameters of the calibration strategy network in the job shop agent; Step 42: using the client agent agent of step 2, extracting feature representations of each process node for each client agent's client agent agent; Step 43: The bidding strategy network of each client agent's respective client agent agent uses the final node feature of each process in the set of to-be-processed processes at the current decision moment as the input of the bidding strategy network, and the bidding strategy network outputs the respective bidding process set to obtain the bidding process set bid of all client agent agents. set Specifically, the bidding process sets of N customer agent agents are summarized to obtain the bidding process set bid of all customer agent agents set , which can be expressed as: Step 44: The job shop agent receives the bid process set bid from all customer agent agents set , and the final node characteristics of each process, and the calibration strategy network is used to execute the autoregressive calibration decision to obtain the calibration decision result; wherein, during the execution of the autoregressive calibration decision, when the calibration sequence length corresponding to the autoregressive calibration decision reaches the number of machines m or all bidding processes have been scheduled, the output of the autoregressive calibration decision is stopped for the calibration strategy network; Step 45, executing the calibration decision result obtained in step 44, and updating the multi-agent job shop scheduling environment based on the job shop agent intelligent agent; Step 46: Calculate the reward value of each customer agent corresponding to the target type according to the target type of each customer agent. The calculation formula is: In formula (14), Represents the reward value, The client agent CA of the i-th client i The number of scheduled operations, Indicates the scheduled process node The value of; wherein the state value network of the client agent intelligent body outputs the value of the scheduled process node; Step 47: Based on the value of the process node, the reward value of the multi-agent job shop is calculated, and the calculation formula is: In formula (15), represents the reward value of the multi-agent job shop, N represents the number of customer agents, The client agent CA of the i-th client i The number of scheduled operations, Indicates the scheduled process node the value of Step 48: Repeat steps 42 to 47 until all processes of each customer agent are scheduled; Step 49: Use a multi-agent deep deterministic policy gradient algorithm to train the customer agent and the job shop agent; Step 410, repeat steps 42 to 49 until the strategies of the customer agent and the job shop agent converge.
Citation Information
Patent Citations
Parallel machine multi-agent auction negotiation scheduling method
CN112101712A
Flexible job shop scheduling method based on deep reinforcement learning and multi-agent graph
CN113361915A
Single-piece job shop scheduling method based on Deep Q-network deep reinforcement learning
CN113792924A
Micro-service-multi-agent factory scheduling model based on deep reinforcement learning network
CN115437321A
Equipment manufacturing workshop intelligent scheduling method and system based on deep reinforcement learning
CN116542445A