Low-carbon flexible job shop scheduling method based on deep reinforcement learning

By constructing a low-carbon flexible job shop scheduling method based on deep reinforcement learning, the problems of insufficient adaptability of existing methods in adapting to heterogeneous features and multi-objective optimization are solved, efficient low-carbon scheduling decisions and carbon emission reduction effects are achieved, and the scheduling needs of large-scale dynamic scenarios are adapted.

CN120806758APending Publication Date: 2025-10-17HIGH FASHION CHINA CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510769360.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing low-carbon flexible job shop scheduling methods based on deep reinforcement learning have difficulty adapting to scheduling instances of different sizes and heterogeneous characteristics. Multi-objective optimization lacks adaptability and lacks explicit modeling of energy consumption characteristics, which limits the carbon sensitivity and environmental adaptability of the model.

Method used

A low-carbon flexible job shop scheduling method based on deep reinforcement learning is constructed. By defining the LC-FJSP mathematical model that includes machine operation energy consumption and no-load energy consumption, a disjunctive graph model of the process-machine constraint relationship is established, and a low-carbon graph attention network LC-GAT is constructed. Combined with the Bayesian optimization module for training and tuning, an actor-critic decision network is designed to implement an end-to-end deep reinforcement learning framework LCGRL.

Benefits of technology

It enables efficient learning of low-carbon scheduling decision rules in complex multi-objective scenarios, improving scheduling efficiency and carbon emission reduction effects, adapting to the scheduling needs of large-scale and dynamic uncertain scenarios, and significantly outperforming traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120806758A_ABST
    Figure CN120806758A_ABST
Patent Text Reader

Abstract

The invention relates to a low-carbon flexible job shop scheduling method based on deep reinforcement learning. The method comprises the following steps: S1, defining an LC-FJSP problem and constructing a disjunction graph model; s2, constructing a problem expression model based on a Markov decision process; s3, introducing a low-carbon map attention network, and integrating multiple attention modules; s4, improving the generalization ability of the model by adopting a graph pooling technology; s5, optimizing a solving process in combination with a Bayesian optimization method; s6, deep reinforcement learning training and verification tuning are carried out; and S7, performing job shop scheduling by using the trained model. According to the method, deep reinforcement learning and Bayesian optimization are combined, an efficient workshop scheduling decision in a low-carbon manufacturing scene is achieved, the constructed scheduling model has high environmental adaptability and energy consumption sensitivity, the low-carbon production scheduling requirement of a complex industrial scene can be effectively met, and intelligent decision support is provided for green manufacturing.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of job shop scheduling, and particularly relates to a low-carbon flexible job shop scheduling method based on deep reinforcement learning. BACKGROUND

[0002] Manufacturing carbon emissions account for 31.8% of total industrial emissions, and low-carbon manufacturing has become a key direction for industrial transformation. Traditional job shop scheduling methods mainly focus on minimizing completion time and cost control, ignoring environmental factors such as machine operating energy consumption and idle energy consumption. As a typical NP-Hard problem, the solution quality of the flexible job shop scheduling problem (FJSP) is directly related to the carbon intensity of the manufacturing process. Studies have shown that optimizing scheduling strategies can reduce energy consumption by 18% to 35%, but achieving both economic and environmental optimization in a dynamic production environment remains a challenge.

[0003] In recent years, deep reinforcement learning (DRL) has shown significant advantages in dynamic scheduling, enabling adaptive decision-making through continuous interaction between agents and the environment, and has become a new technology path for solving complex scheduling problems. However, existing DRL-based methods still have three limitations in the low-carbon scenario: first, the neural network architecture is difficult to adapt to scheduling instances of different scales and heterogeneous characteristics; second, multi-objective optimization relies on artificial weight setting and lacks adaptability; third, there is a lack of explicit modeling of energy consumption characteristics, limiting the model's carbon sensitivity and environmental adaptability. Therefore, there is an urgent need for an intelligent optimization method that combines scheduling efficiency and carbon reduction effect to support green manufacturing practices. SUMMARY

[0004] The present application mainly solves the above problems and provides a low-carbon flexible job shop scheduling method based on deep reinforcement learning with strong environmental adaptability and high energy consumption perception ability.

[0005] The technical solution adopted by the present application to solve its technical problems is a low-carbon flexible job shop scheduling method based on deep reinforcement learning, comprising the following steps:

[0006] S1: define an LC-FJSP mathematical model containing machine operating energy consumption and idle energy consumption and establish a disjunctive graph model of process-machine constraint relationships;

[0007] S2: construct a problem representation model based on Markov decision process, and convert the LC-FJSP mathematical model into a Markov decision process;

[0008] S3: construct a low-carbon graph attention network LC-GAT, including a process feature dynamic perception module, a machine energy consumption competition modeling module, and a multi-granularity feature aggregation module;

[0009] S4: construct an actor-critic decision network;

[0010] S5: configuring a Bayesian optimization module;

[0011] S6: performing deep reinforcement learning training and dynamically refreshing the training set;

[0012] S7: implementing online verification and optimization;

[0013] S8: using the trained model to perform job shop scheduling.

[0014] As a preferred scheme of the above scheme, the LC-FJSP mathematical model in step S1 introduces a carbon emission constraint, and the LC-FJSP model is as follows:

[0015] min f = αC max + βE T

[0016] Wherein, C max is the maximum completion time, E T is the total carbon emission, including carbon emission under processing and carbon emission under idle condition.

[0017] As a preferred scheme of the above scheme, in step S2, the state of the process and the machine is regarded as the overall environment state, and the state of all processes and machines at decision step t constitutes a state s t , the initial state is an LC-FJSP instance, denoted as s0; the process selection and machine allocation are combined into one decision to select an action, and the set of all compatible process-machine pairs is defined as the action space; as time goes on, more and more processes are arranged for processing, and the action space becomes smaller and smaller; at decision step t, the agent samples in the action space under state s t , and takes action a t After that, it transitions to the next new state s t+1 ; the reward function r t at stage t is defined as f(s t )-f(s t+1 ), wherein f represents the value of αC t (s max )+ βE t (s T ) under the current state s t . When the discount factor γ = 1, the reward accumulation of each step can obtain

[0018] As a preferred scheme of the above scheme, the process feature dynamic perception module includes the process node The predecessor O i,j-1 and the successor O i,j+1Local attention field (|p-j|≤1) is constructed, and a relationship coefficient is calculated:

[0019]

[0020] wherein is a trainable weight matrix, and attention calculation is automatically shielded for non-existing process nodes to avoid invalid feature interference, and a normalization coefficient α i,j,p Weighted aggregation of adjacent node features:

[0021]

[0022] As a preferred scheme of the above scheme, the machine energy consumption competition modeling module includes defining machine M k and the competition process set C q of M kq , calculating an energy consumption competition coefficient:

[0023]

[0024] Fusing machine features and competition coefficients to generate attention weights, and obtaining competition-aware attention:

[0025]

[0026] and outputting machine node features through an ELU activation function wherein the input features is a weight matrix, is a linear transformation.

[0027] As a preferred scheme of the above scheme, the multi-granularity feature aggregation module includes setting H independent attention heads, fusing multi-view features through concatenation and average pooling operations, performing multi-attention expansion, and then performing hierarchical aggregation on node features propagated through L layers to complete hierarchical graph pooling:

[0028]

[0029] As a preferred scheme of the above scheme, in the actor-critic decision network, both the actor and the critic use a multi-layer perceptron MLP, and the parameters are represented by θ and π respectively. The actor network first generates a scalar for each selection action a t , and then uses a softmax function to output the required distribution to obtain a random policy; in the actor-critic decision network, all information related to a t is connected into a single vector, and the vector is input into the MLP θ

[0030]

[0031] Select action a t The probability is:

[0032]

[0033] The critic will global features As input, a scalar v(s t ), as an estimate of the state value.

[0034] As a preferred solution of the above scheme, in the Bayesian optimization module, a Gaussian process model is established according to the Bayesian optimization formula to describe the black box function f(C, E), f(C, E) = αC + βE; a point (C) that minimizes the objective function is found under the current Gaussian process model. t+1 , E t+1 ) to select the optimal sampling point, namely:

[0035] (C {t+1} , E {t+1} ) = {argmin} {(C,E)∈X} E[f(C, E)|X, y].

[0036] ; At point (C t+1 , E t+1 ) observation function value f t+1 =f(C t+1 , E t+1 ), then (C t+1 , E t+1 ) and f t+1 Add to the sample points and function values ​​to update the Gaussian process model; use Bayes' theorem and Gaussian process regression to update the mean vector and covariance matrix of the Gaussian process model; repeat the above steps until convergence or the preset number of iterations is reached. The final mean vector μ(X) can be used to estimate the value of μ, specifically expressed as:

[0037] α=μ C / μ,β=μ E / μ

[0038] Among them, μ C and μ E They represent the mean of C and E in the input space, and μ represents the mean of the function values ​​at all points in the input space.

[0039] As a preferred solution of the above scheme, in step S6, a proximal policy optimization algorithm is adopted to interact in parallel on the training set for multiple rounds to dynamically update the network parameters θ, π; and after completing N episodes, the training instance set is resampled to enhance the generalization ability of the model.

[0040] As a preferred solution of the above solution, in the step S7, the model performance is tested on the fixed validation set $X_{val}$ every N episodes, and the weight coefficients a and b are updated by Bayesian optimization.

[0041] The advantages of the present application are: an end-to-end deep reinforcement learning framework LCGRL is constructed, Bayesian optimization is combined with deep reinforcement learning, the state of flexible job shop scheduling is accurately described through a heterogeneous graph model, and the heterogeneous node features are extracted based on a low-carbon graph attention network (LCGAN), which can efficiently learn low-carbon scheduling decision rules and exhibit excellent generalization performance and scheduling efficiency in complex multi-objective scenarios; the multi-objective weight parameter optimization problem is systematically solved, the actor-critic decision network architecture is designed by integrating Bayesian optimization, the weight coefficients of each optimization objective are dynamically adjusted, the reward function convergence speed is significantly improved, and the generated scheduling decision is more in line with the low-carbon goal demand, which has a significant performance advantage over traditional heuristic rules and existing reinforcement learning methods; the practical needs of industrial scenarios are fully considered, the PPO algorithm is used for reinforcement training of the decision network, and the constructed model not only supports joint optimization of operation sequencing and machine allocation, but also can adapt to diversified scheduling requirements on large-scale and public data sets, providing a solid technical foundation for the expansion and application of dynamic uncertainty scenarios. BRIEF DESCRIPTION OF DRAWINGS

[0042] Figure 1 LC-FJSP solving algorithm framework.

[0043] Figure 2 LCGAN network architecture. DETAILED DESCRIPTION

[0044] The technical solutions of the present application will be further described below through embodiments and in combination with the drawings.

[0045] Embodiment:

[0046] The low-carbon flexible job shop scheduling method based on deep reinforcement learning of the present embodiment comprises the following steps:

[0047] S1: defining an LC-FJSP mathematical model containing machine running energy consumption and idle energy consumption and establishing an analytical graph model of process-machine constraint relationship;

[0048] S2: constructing a problem representation model based on Markov decision process, and converting the LC-FJSP mathematical model into a Markov decision process;

[0049] S3: constructing a low-carbon graph attention network LCGAT, including a process feature dynamic perception module, a machine energy consumption competition modeling module and a multi-granularity feature aggregation module;

[0050] S4: Constructing the actor-critic decision network;

[0051] S5: Configure the Bayesian optimization module;

[0052] S6: Perform deep reinforcement learning training and dynamically refresh the training set;

[0053] S7: Implement online verification and tuning;

[0054] S8: Use the trained model for job shop scheduling.

[0055] In step S1, the LC-FJSP model introduces carbon emission constraints based on the traditional FJSP, with the maximum completion time and the minimization of the weighted sum of carbon emissions as the optimization objectives. Carbon emissions under processing conditions:

[0056]

[0057] Carbon emissions when unladen:

[0058]

[0059] Total carbon emissions:

[0060] E T =E1+E2

[0061] The LC-FJSP model is expressed as

[0062] minf=αC max +βE T

[0063] st

[0064]

[0065] Where n is the total number of workpieces; m is the total number of machines; J i represents the i-th workpiece; M k represents the kth machine; n i For workpiece J i The total number of processes included; ij Indicates workpiece J i The jth process; M ij It is process O ij A collection of optional processing machines; ijk For process O ij On machine M k Processing time on S ij and C ij Respectively represent process O ij Processing start and end time; C i For workpiece J itotal completion time; C max maximum completion time; ST k and CT k represent the start and end time of machine M k ; p ijk is the unit time energy consumption of operation O ij on machine M k ; is the idle energy consumption rate of machine M k ; a e is the carbon emission coefficient of electric energy; E T represents the total carbon emission. The decision variable is defined as: x ijk = 1 indicates that operation O ij is processed on machine M k , otherwise 0; y iji′j′,k = 1 indicates that operation O k is processed on machine M ij in preference to operation O' i j' on machine M t , otherwise 0, and L is a sufficiently large positive number. In this embodiment, the LC-FJSP is modeled as a sequential decision problem, and the scheduling process is carried out in an iterative manner: in each decision state, the to-be-scheduled operation is assigned to a compatible machine until all operations are scheduled, realizing efficient solution to complex scheduling tasks, and the overall framework of the algorithm is shown in Figure 1 .

[0066] In step S2, the state, action, state transition and reward of LC-FJSP are first defined, which is converted into a Markov decision process. The operation selection and machine selection are considered as a whole, and a probability distribution is output, and then a greedy algorithm is used to preferentially select the operation-machine pair with the highest score. The LC-FJSP scheduling process can be understood as assigning a ready operation to a compatible idle machine. The scheduling process considered here is as follows: at each decision step t (initial time 0 or operation completion), the agent observes the current system state s t and takes action a t , i.e. assigns the un-planned operation to an idle compatible machine and executes it from the current time T(t). Then, the environment state is transferred to the next decision step t+1. The process is iteratively executed until all operations are scheduled. The corresponding MDP is defined as follows:

[0067] The state feature should be able to describe the main features and changes of the scheduling environment. Here, the states of operations and machines are considered as the overall environment state. The states of all operations and machines at decision step t constitute the state s t, the initial state is an LC-FJSP instance, denoted as s0; process selection and machine allocation are combined into a decision to select an action, and the set of all compatible process-machine pairs is defined as the action space. As time goes by, more and more processes are arranged and processed, and the action space becomes smaller and smaller; at decision step t, in state s t Next, the agent samples in the action space and takes action a t After that, the environment will change and transition to the next new state s t+1 ; By designing a reward function, the agent is guided to select actions that can maximize the completion time of all operations and minimize the total carbon emissions. The reward function r in stage t t is defined as f(s t )-f(s t+1 ), where f represents the current state s t αC under max (s t )+βTCE(s t ) value. When the discount factor γ=1, the cumulative reward of each step can be obtained In a specific problem instance, f(s0) is a constant, which means that minimizing f and maximizing the cumulative reward are equivalent. This embodiment adopts a random strategy π(a t |s t ), for each state s t Define an action set A t The policy distribution is generated by a deep reinforcement learning algorithm that optimizes certain parameters during training to maximize the cumulative reward.

[0068] In step S3, a LCGAN network architecture customized for LC-FJSP is proposed, which obtains the feature representation of process nodes and operation nodes through two attention modules. In the machine feature attention module, the energy consumption feature on the OM arc is added to facilitate the optimization and combination of process features. In order to deal with the weight relationship between time and carbon emissions, the weights of the maximum completion time and carbon emissions are adaptively updated through the Bayesian optimization method to achieve the optimal scheduling solution. The input feature dimensions of the machine and process are d o and d m , the architecture of LCGAN is as follows Figure 2 As shown in the figure, this architecture mainly includes a process feature dynamic perception module, a machine energy consumption competition modeling module and a multi-granularity feature aggregation module.

[0069] Among them, the process feature dynamic perception module includes the process node By Pre-O i,j-1 With the successor O i,j+1Construct a local attention field (|pj|≤1) and calculate the relationship coefficient:

[0070]

[0071] in is a linear transformation, For a trainable weight matrix, since the predecessors or successors of some operations may not exist or will be removed at a certain step, the attention coefficients of these predecessors and successors are dynamically masked. For each input feature Use the softmax function for all e i,j,p Normalize and get the normalized attention coefficient α I,j,p Finally, by transforming the input features and Perform weighted linear combination and connect it with nonlinear activation function σ to obtain the aggregated adjacent node feature vector

[0072]

[0073] The machine energy consumption competition modeling module includes defining the machine M k With M q The competitive process set C kq , calculate the energy consumption competition coefficient:

[0074]

[0075] The machine features and competition coefficients are combined to generate attention weights to obtain competition-aware attention:

[0076]

[0077] Then output the machine node features through the ELU activation function The input features is the weight matrix, is a linear transformation.

[0078] The multi-granularity feature aggregation module includes setting H independent attention heads, so that process O ij and Machine M k The original features are expressed as and After processing by L LCGANs, the features of attention-weighted aggregation are obtained and To be used for subsequent decision-making tasks. Then, multi-view features are fused through splicing and average pooling operations to perform multi-attention expansion, and then the node features propagated through the L layer are hierarchically aggregated to complete the hierarchical graph pooling and obtain the global features.

[0079]

[0080] In step S4, a decision network is designed based on the reinforcement learning framework of the actor and the critic, in which both the actor and the critic use a multi-layer perceptron (MLP) with parameters represented by θ and π, respectively. The actor network first generates a value function for each action a t A scalar is generated, and then a distribution is output using a softmax function, obtaining a random policy for guiding the behavior of the agent in the environment. In this embodiment, a t All relevant information is connected into a single vector, and the vector is input into an MLP θ , as follows:

[0081]

[0082] The probability of selecting an action a t is:

[0083]

[0084] The critic represents a value function network for evaluating the value of taking a certain action in different states, i.e., predicting the cumulative reward under a given policy. It takes global features as input and generates a scalar v(s t ) as an estimate of the state value. The goal of the critic network is to estimate the state value as accurately as possible, thereby providing better feedback and guidance to the actor network to help it improve the policy to obtain higher rewards.

[0085] In step S5, a Bayesian optimization method is used to determine the reward function weight. By selecting appropriate sampling points in the search space and adjusting the positions of the sampling points according to the observed results, the optimal solution is gradually approached. The goal of this embodiment is to optimize the black-box function f(C,E)=αC+βE, where α and β are the coefficients to be optimized, and some sample points (C i ,E i ) and their corresponding function values f i =f(C i ,E i ) are obtained using the results of the decision network. According to the Bayesian optimization formula, a Gaussian process model is established to describe f(C,E).

[0086] To select the optimal sampling point, a point (C t+1 ,E t+1 ) that minimizes the objective function under the current Gaussian process model needs to be found, i.e.:

[0087] (C {t+1} ,E {t+1}) = {argmin} {(C,E)∈X} E[f(C,E)∣X,y].

[0088] To update the Gaussian process model, the function value f t+1 = f(C t+1 ,E t+1 ) needs to be observed at the point (C t+1 ,E t+1 ), and then (C t+1 ,E t+1 ) and f t+1 are added to the sample points and function values. Next, using the method of Bayesian theorem and Gaussian process regression, the mean vector and covariance matrix of the Gaussian process model are updated.

[0089] The above steps are repeatedly performed until convergence or a preset number of iterations is reached. The final mean vector μ(X) can be used to estimate the value of μ. Specifically, it is represented as:

[0090] α = μ C / μ, β = μ E / μ

[0091] where μ C and μ E represent the mean values of C and E in the input space, respectively, and μ represents the mean value of the function value at all points in the input space.

[0092] In step S6, the proximal policy optimization algorithm is used to interact multiple rounds in parallel on the training set, dynamically updating the network parameters θ, π; and after completing N episodes, the training instance set is resampled to enhance the model generalization ability.

[0093] In step S7, the model performance is tested on the fixed validation set $X_{val}$ every N episodes, and the weight coefficients α and β are updated through Bayesian optimization.

[0094] The low-carbon flexible job shop scheduling method based on deep reinforcement learning in the embodiment constructs an end-to-end deep reinforcement learning framework LCGRL, combines Bayesian optimization with deep reinforcement learning, accurately describes the flexible job shop scheduling state through a heterogeneous graph model, extracts heterogeneous node features based on a low-carbon graph attention network (LCGAN), and can efficiently learn low-carbon scheduling decision rules in actual applications, and exhibits excellent generalization performance and scheduling efficiency in complex multi-objective scenarios; systematically solves the multi-objective weight parameter optimization problem, dynamically adjusts the weight coefficients of each optimization objective by designing an actor-critic decision network architecture that integrates Bayesian optimization, significantly improves the reward function convergence speed, and generates scheduling decisions that are more in line with low-carbon target requirements, thereby having significant performance advantages over traditional heuristic rules and existing reinforcement learning methods; fully considers the practical needs of industrial scenarios, uses the PPO algorithm to reinforce the training of the decision network, and the constructed model not only supports joint optimization of operation sequencing and machine allocation, but also can adapt to diversified scheduling requirements on large-scale and public datasets, thereby providing a solid technical foundation for the expansion and application of dynamic uncertainty scenarios.

[0095] The specific embodiments described herein merely exemplify the spirit of the present application. Those skilled in the art of the present application can make various modifications or supplements to the described specific embodiments or use similar ways to replace them, without deviating from the spirit of the present application or exceeding the scope defined by the appended claims.

Claims

1. A low-carbon flexible job shop scheduling method based on deep reinforcement learning, characterized by: The following steps are involved: S1: Define the LC-FJSP mathematical model including machine operation energy consumption and no-load energy consumption and establish a disjunctive graph model of process-machine constraint relationship; S2: Construct a problem formulation model based on the Markov decision process and transform the LC-FJSP mathematical model into a Markov decision process; S3: Build a low-carbon graph attention network (LC-GAT), which includes a process feature dynamic perception module, a machine energy consumption competition modeling module, and a multi-granularity feature aggregation module. S4: Constructing the actor-critic decision network; S5: Configure the Bayesian optimization module; S6: Perform deep reinforcement learning training and dynamically refresh the training set; S7: Implement online verification and tuning; S8: Use the trained model for job shop scheduling.

2. The low-carbon flexible job shop scheduling method based on deep reinforcement learning according to claim 1 is characterized by: In step S1, the LC-FJSP mathematical model introduces carbon emission constraints. The LC-FJSP model is as follows: minf=αC max +βE T Among them, C max is the maximum completion time, E T is the total carbon emissions, including carbon emissions under processing and carbon emissions under no-load conditions.

3. The low-carbon flexible job shop scheduling method based on deep reinforcement learning according to claim 1 is characterized by: In step S2, the states of the processes and machines are regarded as the overall environment state, and the states of all processes and machines at the decision step t constitute the state s. t The initial state is an LC-FJSP instance, denoted as s0; the process selection and machine allocation are combined into a decision to select the action, and the set of all compatible process-machine pairs is defined as the action space. As time goes by, more and more processes are arranged and processed, and the action space becomes smaller and smaller; at decision step t, in state s t Next, the agent samples in the action space and takes action a t After that, transition to the next new state s t+1 ; The reward function r of stage t t is defined as f(s t )-f(s t+1 ), where f represents the current state s t αC under max (s t )+βE T (s t ) value. When the discount factor γ=1, the cumulative reward of each step can be obtained 4. The low-carbon flexible job shop scheduling method based on deep reinforcement learning according to claim 1 is characterized by: The process feature dynamic perception module includes the process node By Pre-O i,j-1 With the successor O i,j+1 Construct a local attention field (|pj|≤1) and calculate the relationship coefficient: in is a linear transformation, It is a trainable weight matrix that automatically blocks attention calculation for non-existent process nodes to avoid invalid feature interference, and then uses the normalization coefficient α i,j,p Weighted aggregation of adjacent node features: 。 5. The low-carbon flexible job shop scheduling method based on deep reinforcement learning according to claim 1 is characterized by: The machine energy consumption competition modeling module includes defining the machine M k With M q The competitive process set C kq , calculate the energy consumption competition coefficient: The machine features and competition coefficients are combined to generate attention weights to obtain competition-aware attention: Then output the machine node features through the ELU activation function The input features is the weight matrix, is a linear transformation.

6. The low-carbon flexible job shop scheduling method based on deep reinforcement learning according to claim 1 is characterized by: The multi-granularity feature aggregation module includes setting H independent attention heads, fusing multi-view features through splicing and average pooling operations, performing multi-attention expansion, and then hierarchically aggregating node features that have been propagated through L layers to complete hierarchical graph pooling: in, and They are the process O obtained after L LCGAN processing ij and Machine M k characteristics.

7. The low-carbon flexible job shop scheduling method based on deep reinforcement learning according to claim 1 is characterized by: In the actor-critic decision network, both the actor and the critic use multi-layer perceptrons (MLPs), with parameters θ and π, respectively. The actor network first selects an action a for each t Generate a scalar, then use the softmax function to output the required distribution to obtain a random strategy; in the actor-critic decision network, a t All relevant information is concatenated into a single vector and fed into the MLP θ Select action a t The probability is: The critic will global features As input, a scalar v(s t ), as an estimate of the state value.

8. The low-carbon flexible job shop scheduling method based on deep reinforcement learning according to claim 1 is characterized by: In the Bayesian optimization module, a Gaussian process model is established according to the Bayesian optimization formula to describe the black box function f(C, E), f(C, E) = αC + βE; a point (C) that minimizes the objective function is found under the current Gaussian process model. t+1 ,E t+1 ) to select the optimal sampling point, namely: (C {t+1} ,ITS {t+1} )={argmin} {(C,E)∈X} E[f(C,E)∣X,y]. At point (C t+1 ,E t+1 ) observation function value f t+1 =f(C t+1 ,E t+1 ), then (C t+1 ,E t+1 ) and f t+1 Add to the sample points and function values ​​to update the Gaussian process model; use Bayes' theorem and Gaussian process regression to update the mean vector and covariance matrix of the Gaussian process model; repeat the above steps until convergence or the preset number of iterations is reached. The final mean vector μ(X) can be used to estimate the value of μ, specifically expressed as: a=m C / μ, β=μ E / m Among them, μ C and μ E They represent the mean of C and E in the input space, and μ represents the mean of the function values ​​at all points in the input space.

9. The low-carbon flexible job shop scheduling method based on deep reinforcement learning according to claim 1 is characterized by: In step S6, a proximal policy optimization algorithm is used to perform multiple rounds of parallel interaction on the training set to dynamically update the network parameters θ, π; and after completing N episodes, the training instance set is resampled to enhance the generalization ability of the model.

10. The low-carbon flexible job shop scheduling method based on deep reinforcement learning according to claim 1 is characterized by: In step S7, the model performance is tested on a fixed validation set $X_{val}$ every N episodes, and the weight coefficients α and β are updated through Bayesian optimization.

Citation Information

Cited By

  • Method, medium and device for solving flexible job-shop scheduling based on deep reinforcement learning

    CN121578760A