Training Method for Flexible Job Shop Scheduling Policy Based on Deep Reinforcement Learning
The flexible work workshop scheduling strategy is trained through the SDAC algorithm model, which solves the problems of instability in training and high resource consumption in the existing technology, and realizes efficient and stable scheduling strategy training and solution.
Patent Information
- Application Number
- CN202311118312.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-31
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-08-31
AI Technical Summary
In the prior art, DQN and PPO algorithms have problems such as unstable training, difficulty convergence, large computing resource consumption, poor generalization and low solution efficiency when training on flexible operation workshop scheduling strategies, and PPO algorithms require a large number of tuning hyperparameters.
The SDAC algorithm model is used to build an actor network and a critic network, design process selection policy network and machine allocation policy network, train the policy network by minimizing the objective function and entropy objective function, and combine the target Q function and soft Q function to optimize the strategy training process.
It improves the training efficiency of the scheduling strategy and the performance and stability of the algorithm, enhances the exploration of the strategy, and can efficiently solve flexible work workshop scheduling problems of different scales.
Smart Images

Figure CN117313792B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of flexible job shop scheduling, and particularly to a training method for flexible job shop scheduling strategies based on deep reinforcement learning. Background Art
[0002] The DQN algorithm uses a deep neural network to approximate the Q function. Network training may be affected by problems such as gradient explosion or disappearance. In the prior art, when using the DQN algorithm for training flexible job shop scheduling strategies, problems such as unstable training and difficulty in convergence are likely to occur. In addition, the DQN algorithm uses a deep neural network, which requires a large amount of computing resources during the training process of flexible job shop scheduling strategies. Therefore, the scheduling strategies trained by the DQN algorithm have poor generalization and low solution efficiency when solving flexible job shop scheduling problem instances of different scales. Additionally, in the prior art, the PPO algorithm is relatively stable in training and has better convergence compared to the DQN algorithm when training flexible job shop scheduling strategies. However, there are multiple hyperparameters in the PPO algorithm that affect the performance and stability of the algorithm, and adjusting these hyperparameters requires a large number of experiments and optimizations. In addition, the scheduling strategies trained by the PPO algorithm are very unstable when solving flexible job shop scheduling problem instances of different scales.
[0003] The invention with the application number CN202310433751.1 discloses a method for multi-objective dynamic flexible job shop scheduling based on deep reinforcement learning regarding energy consumption. The method constructs a high-level deep reinforcement learning network and a low-level deep reinforcement learning network of the double DQN algorithm. The high-level network can control the decisions of the low-level network; input the high-level optimization objectives and state features into the low-level network and optimize the parameters in the network; select appropriate workpieces and processing machines according to the obtained optimization objective values to make the final scheduling optimization objective meet the requirements. The DQN algorithm is adopted in this invention method. The DQN algorithm uses a deep neural network. The scheduling strategies trained by the DQN algorithm of this invention have poor generalization and low solution efficiency when solving flexible job shop scheduling problem instances of different scales. Summary of the Invention
[0004] In view of the above existing technical problems, the present invention provides a training method for flexible job shop scheduling strategies based on deep reinforcement learning.
[0005] The present invention adopts the following specific technical solutions:
[0006] A training method for flexible job shop scheduling strategies based on deep reinforcement learning. The method adopts the SDAC algorithm model and conducts iterative training. The steps of the method are as follows:
[0007] S1: Construct an SDAC algorithm model, where the SDAC algorithm model includes an actor network, a critic network, an entropy objective function, and a sample storage pool D;
[0008] S2: Design a process selection strategy network and a machine allocation strategy network in the actor network to train the process selection strategy and the machine allocation strategy respectively, and train the actor network in the actor network model by minimizing the objective function J π (φ I ).
[0009] The objective function J π (φ I ) has the following formula expression:
[0010]
[0011] where I ∈ {o, m}, o represents the process, m represents the machine, φ I represents the network parameters of the policy π, ∈ t is the input noise vector; ∈ t is sampled from a fixed distribution , D represents the sample storage pool, whose function is to store and reuse the training samples used in the scheduling environment, a t represents the complete action executed at time step t, s t represents the flexible job shop scheduling environment state at time step t;
[0012] S3: Use the entropy objective function to balance the relative importance of the action a and the reward r to control the randomness of the optimal policy;
[0013] S4: Use the target Q function and the soft Q function designed in the critic network to calculate the Q value of a certain state-action pair under the current scheduling policy, and update the critic network by minimizing the target loss function.
[0014] The target loss function J Q (θ) has the following formula expression:
[0015]
[0016] θ represents the parameters of the soft Q network, which are obtained through the training of the soft Q network. Q θ (s t , a t ) is the state-action value function, which is used to obtain the Q value. is the state-action value function approximated according to the parameter θ, which is used to optimize the target loss function J Q (θ) to approximate the state-action value function Q θ (s t , at );
[0017] S5: Use the critic network to control the training of the actor network. After multiple iterations of training, output the final process selection policy network and machine allocation policy network.
[0018] Further, the steps of the iterative training are as follows:
[0019] S1.1: Input the training parameters of the SDAC algorithm model, and initialize the target Q-network parameters in the critic network and the soft Q-network parameters θ, and the policy network parameters φ in the actor network I , and the sample storage pool D;
[0020] S1.2: Before the start of iterative training, reset the flexible job shop scheduling environment state to s t , and during the iterative training process, solve an instance of the flexible job shop scheduling problem to train the scheduling policy;
[0021] S1.3: At time step t, use the process selection policy network and the machine allocation policy network to select actions and Obtain the immediate reward r by executing the complete action a t and transfer to the next state s t ; t+1 ;
[0022] S1.4: The current state transfers from s t to the next state s t+1 , and store the sample (s t , a t , r t , s t+1 ) obtained from the state transfer into the sample storage pool D;
[0023] S1.5: Determine whether the current state s t is the final state. If the current state s t is not the final state, continue to execute step S1.5. Otherwise, randomly extract B samples from the sample storage pool D for learning the scheduling policy;
[0024] S1.6: Update the parameters θ of the soft Q-network, the parameters φ of the policy network I and the entropy parameter α;
[0025] S1.7: Determine whether B samples have been learned. If all B samples have been learned, execute step S1.8. Otherwise, jump to step S1.7;
[0026] S1.8: Determine whether all K flexible job shop scheduling instances have been solved. If all K instances have been solved, execute step S1.9; otherwise, jump to step S1.3.
[0027] S1.9: Determine whether the entire training iteration process has reached the maximum number of iterations. If the maximum number of iterations has been reached then execute step S1.10; otherwise, jump to step S1.2.
[0028] S1.10: Output the trained operation selection policy network and machine allocation policy network to generate the final scheduling policy.
[0029] Further, the training parameters input to the SDAC algorithm model in step S1.1 include the maximum number of iterations the entropy parameter α, the discount factor γ, and the soft update weight parameter ζ.
[0030] Further, the soft Q-network parameters in step S1.6 are updated by calculating the gradient of the objective loss function J π (φ I ) of the critic network.
[0031] Further, the update of the entropy parameter α in step S1.6 is through the minimization objective function J(a) of the entropy objective function.
[0032] Further, the parameters of the policy network in step S1.6 are updated by calculating the gradient of the minimization objective function J Q (θ) of the actor network.
[0033] Further, the update of the entropy parameter α in step S1.6 is through the minimization objective function J(α) of the entropy objective function. The formula expression of J(α) is:
[0034]
[0035] where is a constant representing the threshold of the minimum policy entropy.
[0036] In the entropy objective function is the entropy coefficient, which is a hyperparameter that controls the weight of the entropy term in the objective function. A larger entropy coefficient encourages a larger policy entropy, thus promoting more exploration. However, if the entropy coefficient is set too large, it may cause the policy to be too random and affect performance. By adjusting the size of the entropy coefficient, a trade-off can be made between actions and rewards. When the entropy coefficient is small, the influence of the reward term is greater, and actions that are known to obtain high rewards are more likely to be selected. When the entropy coefficient is large, the influence of the entropy term is greater, and a certain degree of randomness is more likely to be maintained in order to better explore the environment.
[0037] Furthermore, the objective loss function J Q (θ) of the critic network is expressed by the formula:
[0038]
[0039] where is the target Q-value estimated by the target Q-network, γ is the discount factor and 0 ≤ γ ≤ 1, represents the probability of state transition.
[0040] Furthermore, the parameters of the target Q-network are updated by soft update. In the initial state and afterwards it is soft-updated through a formula, and the expression of the formula is:
[0041]
[0042] where ζ is a soft update weight parameter.
[0043] Furthermore, the immediate reward r t is the difference in the makespan from the current state s t to the next state s t+1 , that is, r t = r t (s t , a t , s t+1 ) = C max (s t ) - C max (s t+1 ). The cumulative reward is set as where C max (s end ) = C max , and C max is the makespan. The operation selection policy network and the machine selection policy network in the actor network are used to perform the operation selection action and the machine allocation action and form a complete action a t in the current state s t for each execution of a complete action a t an immediate reward r is obtained through the reward function t and a new state s is generated t+1 ; the critic network trains the actor network through the obtained immediate reward r t to make the process selection policy network and the machine selection policy network in the actor network execute appropriate process selection actions and machine allocation actions to maximize the final cumulative reward
[0044] Furthermore, the method uses experiments to verify its feasibility
[0045] The beneficial effects of the present invention are as follows
[0046] The present invention proposes a training method for a flexible job shop scheduling strategy based on deep reinforcement learning. The model uses the SDAC algorithm model as the deep reinforcement learning model. The method of the present invention trains the process selection strategy and the machine allocation strategy by designing a process selection policy network and a machine allocation policy network respectively, improving the efficiency of policy training. In addition, the designed target Q network and soft Q network in the present invention effectively improve the performance of the algorithm, as well as the convergence and stability of the algorithm. Moreover, the designed entropy objective function increases the exploration of actions during the training of the scheduling strategy, which is beneficial to learning more scheduling strategies. In summary, the present invention not only improves the training efficiency of the scheduling strategy, but also the scheduling strategy trained by the present invention can efficiently solve flexible job shop scheduling problem instances of various scales BRIEF DESCRIPTION OF THE DRAWINGS
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts
[0048] Figure 1 is the training flowchart of the SDAC algorithm
[0049] Figure 2 is a simple schematic diagram of the actor network
[0050] Figure 3 is the makespan convergence curve obtained by using the SDAC algorithm for policy training on flexible job shop scheduling problem instances of different scales
[0051] The following further explains and clarifies in conjunction with embodiments, but the specific embodiments do not limit the present invention in any form.
[0052] Embodiment 1
[0053] As Figures 1 to 3 shown, the embodiment of the present invention discloses a training method for a flexible job shop scheduling strategy based on deep reinforcement learning, including the following steps:
[0054] For the training method of the flexible job shop scheduling strategy based on deep reinforcement learning, the method adopts the SDAC algorithm model and performs iterative training, and the steps of the method are as follows:
[0055] S1: Construct the SDAC algorithm model, and the SDAC algorithm model includes an actor network, a critic network, an entropy objective function, and a sample storage pool D;
[0056] S2: Design a process selection strategy network and a machine allocation strategy network in the actor network to train the process selection strategy and the machine allocation strategy respectively, and train the actor network in the actor network model by minimizing the objective function J π (φ I );
[0057] S3: Use the entropy objective function to balance the relative importance of the action a and the reward r to control the randomness of the optimal strategy;
[0058] S4: Use the target Q function and the soft Q function designed in the critic network to calculate the Q value of a certain state-action pair under the current scheduling strategy, and update the critic network by minimizing the objective loss function;
[0059] S5: Use the critic network to control the training of the actor network, and output the final process selection strategy network and machine allocation strategy network after multiple iterative trainings.
[0060] Furthermore, the steps of performing iterative training using the SDAC algorithm model are as follows:
[0061] S1.1: Input the maximum number of iterations the entropy parameter α, the discount factor γ, and the soft update weight parameter ζ, and initialize the target Q network parameters in the critic network and the soft Q network parameters θ, the policy network parameters φ in the actor network I , and the sample storage pool D;
[0062] S1.2: Before the start of iterative training, reset the flexible job shop scheduling environment state to s t, during the iterative training process, an instance of the flexible job shop scheduling problem is taken to solve and used to train the scheduling strategy;
[0063] S1.3: At time step t, use the operation selection policy network and the machine allocation policy network to select actions respectively and Obtain the immediate reward r t by executing the complete action a t and transfer to the next state s t+1 ;
[0064] S1.4: The current state transfers from s t to the next state s t+1 , and store the sample (s t , a t , r t , s t+1 ) obtained from the state transfer into the sample storage pool D;
[0065] S1.5: Determine whether the current state s t is the final state. If the current state s t is not the final state, continue to execute step S1.5. Otherwise, randomly extract B samples from the sample storage pool D for learning the scheduling strategy;
[0066] S1.6: Update the parameters θ of the soft Q network, the parameters φ of the policy network I and the entropy parameter α;
[0067] The soft Q network parameter θ is updated by calculating the gradient of the objective loss function J π (φ I ) in the critic network. The formula for calculating the critic network to minimize the objective function J π (φ I ) is:
[0068]
[0069] where The formula is:
[0070]
[0071] θ represents the parameters of the soft Q network, obtained through training the soft Q network, represents the parameters of the target Q network, obtained through training the target Q network. Q θ (s t , a t ) is the estimated value of the Q-value function, is the target Q-value estimated by the target Q network, γ is the discount factor and 0 ≤ γ ≤ 1, Represents the probability of state transition.
[0072] The entropy parameter α is updated by minimizing the objective function J(α) of the entropy objective function. The calculation formula of J(α) is as follows:
[0073]
[0074] where is a constant representing the threshold of the minimum policy entropy.
[0075] The parameters φ of the policy network I are updated by calculating the gradient J Q (θ) of the minimization objective function in the actor network. The formula expression of J π (φ I ) is:
[0076]
[0077] where I ∈ {o, m}, φ I represents the network parameters of the policy π, and ∈ t is the input noise vector. ∈ t is sampled from the fixed distribution .
[0078] S1.7: Determine whether B samples have been learned. If all B samples have been learned, execute step S1.8; otherwise, jump to step S1.7;
[0079] S1.8: Determine whether all K flexible job shop scheduling instances have been solved. If all K instances have been solved, execute step S1.9; otherwise, jump to step S1.3;
[0080] S1.9: Determine whether the entire training iteration process has reached the maximum number of iterations If the maximum number of iterations has been reached then execute step S1.10; otherwise, jump to step S1.2;
[0081] S1.10: Output the trained operation selection policy network and machine allocation policy network to generate the final scheduling policy.
[0082] Furthermore, the parameters of the target Q network are updated by soft update. In the initial state and afterwards are soft-updated by a formula. The formula expression is:
[0083]
[0084] In formula (5) thereof, ζ is a soft update weight parameter.
[0085] Further, the sample storage pool is used to store and reuse the training samples used in the scheduling environment.
[0086] Further, when using the SDAC algorithm for policy training, six different scales of flexible job shop scheduling (FJSP) instances randomly generated are used for policy training, including 6×6, 10×5, 10×10, 15×15, 20×10, and 30×10, as Figure 3 shown, Figure 3 shows the convergence curves of the makespan obtained by policy training on six different scales of randomly generated FJSP instances. It can be seen from Figure 2 that after about 50 iterations, the fluctuations of the makespan convergence curve become smaller, indicating that using the SDAC algorithm for policy training can converge quickly.
[0087] The beneficial effects of the present invention are as follows:
[0088] The present invention proposes a flexible job shop scheduling policy training method based on deep reinforcement learning. The model uses the SDAC algorithm model as the deep reinforcement learning model. The method of the present invention trains the operation selection policy and the machine allocation policy by designing an operation selection policy network and a machine allocation policy network respectively, improving the efficiency of policy training. In addition, the designed target Q network and soft Q network of the present invention effectively improve the performance of the algorithm, as well as the convergence and stability of the algorithm. Moreover, the designed entropy objective function increases the exploration of actions during scheduling policy training, which is beneficial to learning more scheduling policies. In summary, the present invention not only improves the training efficiency of the scheduling policy, but also the scheduling policy trained by the present invention can efficiently solve flexible job shop scheduling problem instances of various scales.
[0089] Embodiment 2
[0090] On the basis of Embodiment 1, in order to verify the feasibility of the SDAC algorithm model, in this experiment, the SDAC algorithm adopted by the present invention, the DQN algorithm, and the PPO algorithm are applied to the same data set to compare the solution quality of the SDAC algorithm, the DQN algorithm, and the PPO algorithm adopted by the present invention.
[0091] Ten FJSP instances of different scales were extracted from the Behnke dataset. The SDAC algorithm, DQN algorithm, and PPO algorithm were used to solve these 10 instances of different scales respectively. By comparing the solution quality (i.e., the makespan) and the solution time (i.e., the running time of running these three algorithms on the same device) respectively, the superiority of the proposed SDAC algorithm was highlighted. As shown in Table 1, Table 1 shows the comparison of the solution quality of the SDAC algorithm, DQN algorithm, and PPO algorithm on the above-mentioned extracted Behnke dataset. In Table 1, C max is the makespan. The smaller the makespan, the better the solution quality. UB is the best-known solution (i.e., the smallest makespan known for this FJSP instance), and Gap is the performance evaluation index. The Gap value is calculated by the following formula:
[0092]
[0093] As can be seen from Table 1 below, the makespan C obtained by using the SDAC algorithm on all FJSP instances max is much smaller than the makespan C obtained by using the DQN algorithm and PPO algorithm max , indicating that the solution quality obtained by using the SDAC algorithm on all FJSP instances is much better than that of the DQN algorithm and PPO algorithm. The Gap value of the SDAC algorithm on all FJSP instances is much smaller than that of the DQN algorithm and PPO algorithm, indicating that the scheduling performance of the SDAC algorithm on all FJSP instances is much better than that of the DQN algorithm and PPO algorithm. The results in Table 1 show that the method for training the flexible job shop scheduling strategy proposed in the present invention effectively improves the solution quality of solving flexible job shop scheduling problem instances of different scales.
[0094]
[0095] Table 1
[0096] Example 3
[0097] On the basis of Example 1, in order to verify the feasibility of the SDAC algorithm model, in this experiment, the SDAC algorithm adopted in the present invention, the DQN algorithm, and the PPO algorithm were applied to the same dataset, and the solution times of the SDAC algorithm, the DQN algorithm, and the PPO algorithm adopted in the present invention were compared.
[0098] Table 2 lists the solution times (i.e., the running times of the algorithms when solving the FJSP instances) of using the SDAC algorithm, DQN algorithm, and PPO algorithm respectively to solve 10 FJSP instances of different scales on the same device. The shorter the solution time, the faster the solution speed of the algorithm. As can be seen from Table 1, the solution times of the SDAC algorithm are less than those of the DQN algorithm and the PPO algorithm on all FJSP instances, indicating that the method proposed in the present invention for training the flexible job shop scheduling strategy effectively reduces the solution times of solving flexible job shop scheduling problem instances of different scales.
[0099]
[0100]
[0101] Table 2
[0102] Obviously, the above-mentioned embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A training method for flexible job shop scheduling strategy based on deep reinforcement learning, characterized in that The steps of the method are as follows: S1: Construct an SDAC algorithm model, where the SDAC algorithm model includes an actor network, a critic network, an entropy objective function, and a sample storage pool ; S2: Design a process selection strategy network and a machine allocation strategy network in the actor network to train the process selection strategy and the machine allocation strategy respectively, and train the actor network in the actor network model by minimizing the objective function to train the actor network in the actor network model; The minimization objective function in the actor network The formula expression of which is as follows: Among them , represents a process, represents a machine, represents the policy of network parameters, is the input noise vector; is sampled from a fixed distribution , represents a sample storage pool, whose function is to store and reuse the training samples used in the scheduling environment, represents the execution of a complete action at time step , represents the state of the flexible job shop scheduling environment at time step . S3: Balance actions using an entropy objective function and rewards to control the randomness of the optimal policy; S4: Calculate the Q value of a state-action pair under the current scheduling policy by using the target Q function and the soft Q function designed in the critic network, and update the critic network by minimizing the target loss function; The target loss function of the critic network The formula expression is as follows: represent the parameters of the soft Q-network, obtained through training the soft Q-network is the state-action value function, used to obtain the Q value is the state-action value function approximated according to the parameter for optimizing the objective loss function to approximate the state-action value function ; S5: Use the critic network to control the training of the actor network, and output the final process selection policy network and machine allocation policy network after multiple iterative trainings.
2. The training method of the flexible job shop scheduling strategy based on deep reinforcement learning according to claim 1, characterized in that The steps of iterative training using the constructed SDAC algorithm model are as follows: S1.1: Input the training parameters of the SDAC algorithm model, and initialize the target Q-network parameters in the critic network and the soft Q-network parameters , the policy network parameters in the actor network , and the sample storage pool ; S1.2: Before the start of iterative training, reset the flexible job shop scheduling environment state to , and during the iterative training process, take instances of the flexible job shop scheduling problem for solution to train the scheduling strategy; S1.3: At time step , use the process selection policy network and the machine allocation policy network to separately select the action and the machine allocation action . Obtain the immediate reward by executing the complete action and transfer to the next state ; S1.4: The current state transfers from to the next state , and stores the sample obtained from the state transfer in the sample storage pool ; S1.5: Determine the current status Whether it is the final status. If the current status is not the final status, continue to execute step S1.
5. Otherwise, randomly select from the sample storage pool samples for learning the scheduling strategy; S1.6: Update the parameters of the soft Q-network , the parameters of the policy network and the entropy parameter ; S1.7: Determine whether samples have been learned. If all samples have been learned, execute step S1.8; otherwise, jump to step S1.
6. S1.8: Determine if all flexible job shop scheduling instances have been solved. If all instances have been solved, then execute step S1.9; otherwise, jump to step S1.
3. S1.9: Determine whether the maximum number of iterations has been reached in the entire training iteration process If the maximum number of iterations has been reached then execute step S1.10; otherwise, jump to step S1.2 S1.10: Output the trained process selection policy network and machine allocation policy network to generate the final scheduling policy.
3. The training method of the flexible job shop scheduling strategy based on deep reinforcement learning according to claim 2, characterized in that, The training parameters input into the SDAC algorithm model in the step S1.1 include the maximum number of iterations , the entropy parameter , the discount factor , and the soft update weight parameter .
4. The training method for the flexible job shop scheduling strategy based on deep reinforcement learning according to claim 2, characterized in that The soft Q-network parameters in step S1.6 are updated by calculating the gradient of the target loss function in the critic network .
5. The training method of the flexible job shop scheduling strategy based on deep reinforcement learning according to claim 2, wherein The parameters of the policy network in step S1.6 are updated by calculating the gradient of the minimized objective function in the actor network .
6. The training method of the flexible job shop scheduling strategy based on deep reinforcement learning according to claim 2, wherein The entropy parameter in step S1.6 is updated by minimizing the objective function of the entropy objective function , and the has the following formula expression: where is a constant representing the threshold of the minimum policy entropy.
7. The training method of the flexible job shop scheduling strategy based on deep reinforcement learning according to claim 1, characterized in that The target loss function of the critic network in the formula expression is as follows: where is the target Q-value estimated by the target Q-network, is the discount factor and , represents the probability of state transition.
8. The training method of the flexible job shop scheduling strategy based on deep reinforcement learning according to claim 2, wherein The parameters of the target Q network are updated in a soft update manner. In the initial state , and afterwards soft update is performed through a formula, and the expression of the formula is: Among them is a soft update weight parameter.
9. The training method of the flexible job shop scheduling strategy based on deep reinforcement learning according to claim 8, characterized in that The instant reward in step S1.3 has the following formula expression: Among them, the immediate reward is the difference in the maximum completion time for transferring from the current state to the next state, and the maximum completion time is the maximum completion time.
10. The training method of the flexible job shop scheduling strategy based on deep reinforcement learning according to claim 1, characterized in that, The method is verified by experiments.
Citation Information
Patent Citations
Multi-target dynamic flexible job shop scheduling method about energy consumption based on deep reinforcement learning
CN116644902A
Distributed training method and device for deep convolutional neural network, and storage medium
CN112712171A
KR1024407540000B1