An optimization method and device for distributed re-entrant job-shop dynamic scheduling problem

By employing reinforcement learning methods and an improved double-Q learning algorithm in steel enterprises, the scheduling of distributed reentrant workshops was optimized, solving the real-time scheduling problem in dynamic environments, achieving rapid response and efficient resource utilization, and improving production management efficiency and economic benefits.

CN115271346BActive Publication Date: 2026-01-30UNIV OF SCI & TECH BEIJING
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210711419.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-22
Publication Date
2026-01-30
Estimated Expiration
2042-06-22

AI Technical Summary

Technical Problem

In steel enterprises with multi-factory collaborative production, how can we achieve real-time scheduling decisions in the workshop under dynamic conditions, optimize resource utilization, respond quickly to dynamic events, and avoid irrationality caused by manual decision-making and lag in production management?

Method used

By employing reinforcement learning, and by setting state feature vectors, action sets, and reward functions, combined with an improved double-Q learning algorithm, the scheduling scheme of the distributed reentrant workshop is optimized. Electronic devices are used to realize the state transition and action selection of the agent, enabling real-time scheduling.

Benefits of technology

It improved production management efficiency, avoided the irrationality and lag of manual scheduling, enhanced the company's production management capabilities, and improved productivity and economic benefits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115271346B_ABST
    Figure CN115271346B_ABST
Patent Text Reader

Abstract

This invention discloses an optimization method and apparatus for the distributed reentrant shop floor dynamic scheduling problem, relating to the field of shop floor scheduling technology. It includes: acquiring processing information of the shop to be scheduled; setting a state feature vector, action set, and reward function for reinforcement learning; and obtaining a shop floor scheduling scheme based on the processing information, state feature vector, action set, reward function, and an improved double-Q learning algorithm. This invention can solve the problem of how to rationally formulate dynamic scheduling decision-making methods in a distributed reentrant shop floor to achieve coordinated production of three work sections and fully utilize existing resources. This invention can effectively solve the distributed reentrant shop floor dynamic scheduling problem and respond promptly to dynamic disturbance events.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of workshop scheduling technology, and in particular to an optimization method and apparatus for the dynamic scheduling problem of distributed reentrant workshops. Background Technology

[0002] The development of manufacturing and the deepening of economic globalization have promoted the transformation and upgrading of the steel industry. Intense competition has made the traditional single-factory production model in the steel industry unable to meet the needs of enterprises to increase output and enhance their competitiveness. Therefore, enterprises need to transform from single-factory production to multi-factory collaborative production.

[0003] As a crucial component of my country's manufacturing sector, the steel industry has faced severe challenges in recent years. On the one hand, the global steel industry is experiencing overcapacity; on the other hand, as a key target of environmental protection and rectification efforts in my country, the low efficiency and severe losses of small and medium-sized steel enterprises have not been fundamentally improved, and some even face the risk of closure. This has led to increasingly fierce competition within the industry, prompting steel companies to explore ways to improve production efficiency in order to achieve higher economic benefits. Therefore, optimizing the scheduling of workshop production in steel enterprises is an urgent need for many companies and a crucial way to seek a rational production model, improve productivity, increase profits, and survive in the competition.

[0004] Among the diverse range of steel products, seamless steel pipes are widely used and often referred to as the "blood vessels of industry," playing a crucial role in defense, aerospace, and the petroleum industry. They are also essential components in engineering construction and daily life, commonly found in automobile driveshafts, bicycle frames, and steel scaffolding. Due to their wide applications, excellent performance, and precision, seamless steel pipes have become a benchmark for a country's steel pipe technology development. However, deformation conditions and metal strength limit the cold drawing process. To achieve the desired diameter reduction, multiple cold drawing processes are required, necessitating re-entry into the system, increasing machine wear and the likelihood of machine failure. Therefore, incorporating dynamic interference factors into the research scope can prevent irrationalities caused by the workshop's inability to make real-time scheduling and production decisions during dynamic events, thereby strengthening enterprise production management. Summary of the Invention

[0005] This invention addresses the problem of how to rationally formulate real-time scheduling decision-making methods in a dynamic environment for distributed reentrant workshops, so as to achieve rapid response to dynamic events and make full use of existing resources.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0007] On one hand, this invention provides an optimization method for the distributed reentrant shop floor dynamic scheduling problem, implemented by electronic devices, and comprising:

[0008] S1. Obtain the processing information of the workshop to be scheduled; the processing information includes the number of workpieces to be processed, processing time, number of parallel machines, allocation result and delivery time; among which, the processing time includes the processing time of the first section, the processing time of the second section and the processing time of the third section; the allocation result is the result of the workpieces being allocated to different factories in the third section.

[0009] S2. Define the state feature vector, action set, and reward function for reinforcement learning.

[0010] S3. Based on the processing information, state feature vector, action set, reward function, and improved double-Q learning algorithm, the workshop scheduling scheme is obtained.

[0011] Optionally, the set state feature vector in S2 includes the set state feature vector and the empty state.

[0012] The state feature vectors include state feature vector 1, state feature vector 2, state feature vector 3, state feature vector 4, state feature vector 5, and state feature vector 6.

[0013] Among them, the state feature vector 1 represents the ratio of the maximum remaining processing time of the workpiece in the current buffer queue to the average remaining processing time of all workpieces.

[0014] State feature vector 2 represents the ratio of the minimum remaining processing time of the workpiece in the buffer queue to the average remaining processing time of all workpieces.

[0015] State feature vector 3 represents the average waiting time of all artifacts in the buffer queue.

[0016] State feature vector 4 represents the mean relaxation time of all workpieces in the buffer queue.

[0017] The state feature vector 5 represents the ratio of the maximum to the average processing time of the current process in the buffer queue.

[0018] The state feature vector 6 represents the ratio of the minimum to the average processing time of the current process in the buffer queue.

[0019] The empty state includes empty state 1 and empty state 2.

[0020] Among them, empty state 1 indicates that the machine buffer is empty.

[0021] An empty state 2 indicates that there is one and only one workpiece in the machine buffer.

[0022] Optionally, the action set in S2 includes: minimum relaxation time rule (SLT), high response ratio priority rule (HRN), shortest remaining processing time rule (LWR), shortest processing time priority rule (SPT), longest processing time priority rule (LPT), and earliest delivery date rule (EDD).

[0023] The SLT rule means selecting the workpiece with the shortest relaxation time. If the relaxation times are the same, then one of the workpieces with the same relaxation time is randomly selected.

[0024] The HRN rule means selecting the workpiece with the largest ratio of the sum of waiting time and processing time to the processing time. If the ratios are the same, then one of the workpieces with the same ratio is randomly selected.

[0025] The LWR rule means selecting the workpiece with the shortest remaining processing time; if the remaining processing times are the same, then the workpiece with the earliest arrival time is selected.

[0026] The SPT rule selects the workpiece with the shortest processing time in the current process. If the processing times in the current processes are the same, the workpiece with the earliest arrival time is selected.

[0027] The LPT rule means selecting the workpiece with the longest processing time in the current process. If the processing times in the current processes are the same, then the workpiece with the earliest arrival time is selected.

[0028] The EDD rule means selecting the workpiece with the earliest delivery date; if the delivery dates are the same, then the workpiece with the earliest arrival time is selected.

[0029] Optionally, the shop floor scheduling scheme obtained in S3 based on processing information, state feature vectors, action sets, reward functions, and the improved double-Q learning algorithm includes:

[0030] S301. Set the parameters of the improved double-Q learning algorithm; the parameters include the learning rate α, discount factor γ, number of clusters m, and number of training iterations. t The iteration count is iter, the experience sharing rate is K, and the greedy strategy parameter is e.

[0031] S302, Obtain the state of the agent.

[0032] S303, Initialize each agent j Table and Initialize the common Q-value table; let k0 = 0.99, i = 1.

[0033] S304. If i < iter*0.2, then execute S305; if i ≥ iter*0.2, then execute S311.

[0034] S305. Calculate the value of the state feature vector of each agent j based on the agent's state, cluster the values ​​of the state feature vectors of each agent j to obtain the clustering state of the agents; obtain the common Q value based on the clustering state; and use the value Q from the common Q value table with a probability of 1-K. C (s t ,a * Replace the current agent j The values ​​in the table and The values ​​in the table are used to obtain the values ​​of the replaced agent j. Table and surface.

[0035] S306, according to the replacement surface, The table shows the actions selected by the greedy strategy.

[0036] S307. Execute the selected action to transition the agent's state to state s. t+1 And calculate the immediate reward r based on the reward function.

[0037] S308, Update the replacement agent j selected in S306 Table or The table is updated, and the public Q-value table is updated.

[0038] S309, Let K = 0.99 i *k0; Updates the state of agent j.

[0039] S310. Determine whether all workpieces have been scheduled. If they have been scheduled, set i = i + 1 and proceed to S304. If they have not been scheduled, proceed to S305.

[0040] S311, Random selection Table replacement Table or selection Table replacement The table is obtained after assimilation.

[0041] S312. Calculate the value of the state feature vector of each agent, cluster the values ​​of the state feature vectors of each agent to obtain the cluster state of the agents; obtain the common Q value based on the cluster state, and adopt the value of the common Q value table with a probability of 1-K. replace The value of the table.

[0042] S313, Replace with The table selects the maximum value based on the greedy strategy. action Choose an action with a probability of 1-e Or randomly select action a with probability e. t , thus obtaining the action to choose.

[0043] S314. Execute the selected action to transition the agent's state to state s. t+1 And calculate the immediate reward r based on the reward function.

[0044] S315, Updated and Replaced Table and common Q-value table; let K = 0.99 i *k0; Updates the state of the agent.

[0045] S316. Determine whether all workpieces have been scheduled. If they have been scheduled, proceed to S317. If they have not been scheduled, proceed to S312.

[0046] S317. If i < iter, then let i = i + 1 and proceed to execute S312; if i ≥ iter, then output the workshop scheduling scheme obtained through iteration.

[0047] Optionally, in S306, according to the replacement surface, The table and the actions selected by the greedy strategy include:

[0048] Choose an action with a probability of 1-e Or randomly select action a with probability e. t , thus obtaining the action to choose.

[0049] Among them, actions The selection process includes:

[0050] Randomly select the replacement Table or The table serves as the basis for action selection.

[0051] If you select the replacement If the table is used as the basis for action selection, then the replacement will be selected according to the greedy strategy. Maximum value of the table Action as action

[0052] If you select the replacement If the table is used as the basis for action selection, then the replacement will be selected according to the greedy strategy. Maximum value of the table Action as action

[0053] Optionally, the update in S308 is the replacement agent j selected in S306. The table includes:

[0054] When agent j is in a non-empty state, the replaced agent j is updated according to the following equations (1) and (2). surface:

[0055]

[0056]

[0057] Where t represents time s t The state at time t; a t The action at time t; Choose action a for agent j at time t. t The corresponding Q A The table shows the values; α is the learning rate; r t γ is the instant reward at time t; γ is the discount factor; s t+1 The state at time t+1; a t+1 For s t+1 Actions in a certain state; For s t+1 Q of agent j in state A The maximum value in the table; Let s be the time t+1. t+1 In the given state, agent j selects action a. t+1 The corresponding Q B The values ​​in the table.

[0058] When agent j is in empty state 2, the replaced agent j is updated according to the following formula (3). surface:

[0059]

[0060] Optionally, the updated common Q-value table in S308 includes:

[0061] Public Q-value table and intelligent agent Table and The average value of the table is empirically shared and calculated using the following formulas (4) and (5):

[0062]

[0063]

[0064] Among them, s t+1 The state at time t+1; Indicates in s t+1 Time makes The action with the largest value; m (h) This represents the current number of intelligent agents in the factory.

[0065] Optionally, the updated replacement in S315 The table includes:

[0066] When agent j is in a non-empty state, the replaced data is updated according to the following formula (6). surface:

[0067]

[0068] When agent j is in empty state 2, the replaced state is updated according to the following formula (7). surface:

[0069]

[0070] in, For time t The value of the table; s t The state at time t; a t Let r be the action at time t; α be the learning rate; r be the action at time t; α be the learning t γ is the instant reward at time t; γ is the discount factor; s t+1 The state at time t+1; a t+1 For s t+1 The action selected in the current state; Let s be the time t+1. t+1 In state The maximum value of the action.

[0071] Optionally, the updated common Q-value table in S315 includes:

[0072] Public Q-value table and intelligent agent Experience sharing is performed, calculated using the following formula (8):

[0073]

[0074] Among them, s t+1 The state at time t+1; Indicates in s t+1 Time makes The action with the largest value; m (h) This represents the current number of intelligent agents in the factory.

[0075] On the other hand, the present invention provides an optimization apparatus for the distributed reentrant shop floor dynamic scheduling problem. This apparatus is applied to an optimization method for implementing the distributed reentrant shop floor dynamic scheduling problem. The apparatus includes:

[0076] The acquisition module is used to acquire the processing information of the workshop to be scheduled. The processing information includes the number of workpieces to be processed, processing time, number of parallel machines, allocation result and delivery time. The processing time includes the processing time of the first section, the processing time of the second section and the processing time of the third section. The allocation result is the result of the workpieces being allocated to different factories in the third section.

[0077] The configuration module is used to configure the state feature vectors, action set, and reward function for reinforcement learning.

[0078] The output module is used to obtain the workshop scheduling scheme based on the processing information, state feature vector, action set, reward function and improved double-Q learning algorithm.

[0079] Optionally, the set state feature vector includes the set state feature vector and the empty state.

[0080] The state feature vectors include state feature vector 1, state feature vector 2, state feature vector 3, state feature vector 4, state feature vector 5, and state feature vector 6.

[0081] Among them, the state feature vector 1 represents the ratio of the maximum remaining processing time of the workpiece in the current buffer queue to the average remaining processing time of all workpieces.

[0082] State feature vector 2 represents the ratio of the minimum remaining processing time of the workpiece in the buffer queue to the average remaining processing time of all workpieces.

[0083] State feature vector 3 represents the average waiting time of all artifacts in the buffer queue.

[0084] State feature vector 4 represents the mean relaxation time of all workpieces in the buffer queue.

[0085] The state feature vector 5 represents the ratio of the maximum to the average processing time of the current process in the buffer queue.

[0086] The state feature vector 6 represents the ratio of the minimum to the average processing time of the current process in the buffer queue.

[0087] The empty state includes empty state 1 and empty state 2.

[0088] Among them, empty state 1 indicates that the machine buffer is empty.

[0089] An empty state 2 indicates that there is one and only one workpiece in the machine buffer.

[0090] Optionally, the action set includes: minimum relaxation time rule (SLT), high response ratio priority rule (HRN), shortest remaining processing time rule (LWR), shortest processing time priority rule (SPT), longest processing time priority rule (LPT), and earliest delivery date rule (EDD).

[0091] The SLT rule means selecting the workpiece with the shortest relaxation time. If the relaxation times are the same, then one of the workpieces with the same relaxation time is randomly selected.

[0092] The HRN rule means selecting the workpiece with the largest ratio of the sum of waiting time and processing time to the processing time. If the ratios are the same, then one of the workpieces with the same ratio is randomly selected.

[0093] The LWR rule means selecting the workpiece with the shortest remaining processing time; if the remaining processing times are the same, then the workpiece with the earliest arrival time is selected.

[0094] The SPT rule selects the workpiece with the shortest processing time in the current process. If the processing times in the current processes are the same, the workpiece with the earliest arrival time is selected.

[0095] The LPT rule means selecting the workpiece with the longest processing time in the current process. If the processing times in the current processes are the same, then the workpiece with the earliest arrival time is selected.

[0096] The EDD rule means selecting the workpiece with the earliest delivery date; if the delivery dates are the same, then the workpiece with the earliest arrival time is selected.

[0097] Optionally, the output module is further used for:

[0098] S301. Set the parameters of the improved double-Q learning algorithm; the parameters include the learning rate α, discount factor γ, number of clusters m, and number of training iterations. t The iteration count is iter, the experience sharing rate is K, and the greedy strategy parameter is e.

[0099] S302, Obtain the state of the agent.

[0100] S303, Initialize each agent j Table and Initialize the common Q-value table; let k0 = 0.99, i = 1.

[0101] S304. If i < iter*0.2, then execute S305; if i ≥ iter*0.2, then execute S311.

[0102] S305. Calculate the value of the state feature vector of each agent j based on the agent's state, cluster the values ​​of the state feature vectors of each agent j to obtain the clustering state of the agents; obtain the common Q-value based on the clustering state; and use the value of the common Q-value table with a probability of 1-K. Replace the current agent j The values ​​in the table and The values ​​in the table are used to obtain the values ​​of the replaced agent j. Table and surface.

[0103] S306, according to the replacement surface, The table shows the actions selected by the greedy strategy.

[0104] S307. Execute the selected action to transition the agent's state to state s. t+1 And calculate the immediate reward r based on the reward function.

[0105] S308, Update the replacement agent j selected in S306 Table or The table is updated, and the public Q-value table is updated.

[0106] S309, Let K = 0.99 i *k0; Updates the state of agent j.

[0107] S310. Determine whether all workpieces have been scheduled. If they have been scheduled, set i = i + 1 and proceed to S304. If they have not been scheduled, proceed to S305.

[0108] S311, Random selection Table replacement Table or selection Table replacement The table is obtained after assimilation.

[0109] S312. Calculate the value of the state feature vector of each agent, cluster the values ​​of the state feature vectors of each agent to obtain the cluster state of the agents; obtain the common Q value based on the cluster state, and adopt the value of the common Q value table with a probability of 1-K. replace The value of the table.

[0110] S313, Replace with The table selects the maximum value based on the greedy strategy. action Choose an action with a probability of 1-e Or randomly select action a with probability e. t, thus obtaining the action to choose.

[0111] S314. Execute the selected action to transition the agent's state to state s. t+1 And calculate the immediate reward r based on the reward function.

[0112] S315, Updated and Replaced Table and common Q-value table; let K = 0.99 i *k0; Updates the state of the agent.

[0113] S316. Determine whether all workpieces have been scheduled. If they have been scheduled, proceed to S317. If they have not been scheduled, proceed to S312.

[0114] S317. If i < iter, then let i = i + 1 and proceed to execute S312; if i ≥ iter, then output the workshop scheduling scheme obtained through iteration.

[0115] Optionally, the output module is further used for:

[0116] Choose an action with a probability of 1-e Or randomly select action a with probability e. t , thus obtaining the action to choose.

[0117] Among them, actions The selection process includes:

[0118] Randomly select the replacement Table or The table serves as the basis for action selection.

[0119] If you select the replacement If the table is used as the basis for action selection, then the replacement will be selected according to the greedy strategy. Maximum value of the table Action as action

[0120] If you select the replacement If the table is used as the basis for action selection, then the replacement will be selected according to the greedy strategy. Maximum value of the table Action as action

[0121] Optionally, the output module is further used for:

[0122] When agent j is in a non-empty state, the replaced agent j is updated according to the following equations (1) and (2). surface:

[0123]

[0124]

[0125] Where t represents time s t The state at time t; a t The action at time t; Choose action a for agent j at time t. t The corresponding Q A The table shows the values; α is the learning rate; r t γ is the instant reward at time t; γ is the discount factor; s t+1 The state at time t+1; a t+1 For s t+1 Actions in a certain state; For s t+1 Q of agent j in state A The maximum value in the table; Let s be the time t+1. t+1 In the given state, agent j selects action a. t+1 The corresponding Q B The values ​​in the table.

[0126] When agent j is in empty state 2, the replaced agent j is updated according to the following formula (3). surface:

[0127]

[0128] Optionally, the output module is further used for:

[0129] Public Q-value table and intelligent agent Table and The average value of the table is empirically shared and calculated using the following formulas (4) and (5):

[0130]

[0131]

[0132] Among them, s t+1 The state at time t+1; Indicates in s t+1 Time makes The action with the largest value; m (h) This represents the current number of intelligent agents in the factory.

[0133] Optionally, the output module is further used for:

[0134] When agent j is in a non-empty state, the replaced data is updated according to the following formula (6). surface:

[0135]

[0136] When agent j is in empty state 2, the replaced state is updated according to the following formula (7). surface:

[0137]

[0138] in, For time t The value of the table; s t The state at time t; a t Let r be the action at time t; α be the learning rate; r be the action at time t; α be the learning t γ is the instant reward at time t; γ is the discount factor; s t+1 The state at time t+1; a t+1 For s t+1 The action selected in the current state; Let s be the time t+1. t+1 In state The maximum value of the action.

[0139] Optionally, the output module is further used for:

[0140] Public Q-value table and intelligent agent Experience sharing is performed, calculated using the following formula (8):

[0141]

[0142] Among them, s t+1 The state at time t+1; Indicates in s t+1 Time makes The action with the largest value; m (h) This represents the current number of intelligent agents in the factory.

[0143] On the one hand, an electronic device is provided, comprising a processor and a memory, wherein the memory stores at least one instruction, which is loaded and executed by the processor to implement the aforementioned optimization method for the distributed reentrant workshop dynamic scheduling problem.

[0144] On the one hand, a computer-readable storage medium is provided, wherein at least one instruction is stored in the storage medium, and the at least one instruction is loaded and executed by a processor to implement the above-mentioned optimization method for the distributed reentrant workshop dynamic scheduling problem.

[0145] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following:

[0146] In the above scheme, an optimization model is established for the multi-stage distributed reentrant shop floor dynamic scheduling problem. Considering the efficiency of processing and production, the evaluation indicators for scheduling are set as total delay and completion time. Reinforcement learning is adopted for real-time scheduling, transforming the distributed reentrant shop floor dynamic scheduling problem into a reinforcement learning problem. A reward function is designed based on the two evaluation indicators of completion time and total delay. The classic Q-learning algorithm cannot maintain good performance in complex production environments with multiple agents, multiple work sections, multiple factories, and reentrancy. Therefore, the classic Q-learning algorithm is improved to obtain an improved double-Q learning algorithm. To avoid Q-learning overestimating action values, a double-Q table learning mechanism is designed in the early stage of iteration; and the greedy selection coefficient is updated with a small step size in the early stage to increase the probability of exploration; in the later stage of iteration, a single-Q table learning mechanism is adopted to increase the step size of the greedy selection; the multi-agent experience interaction mechanism runs through the entire process of the algorithm. It aims to respond promptly to dynamic events such as machine failures and random arrival of workpieces in the production process and achieve real-time scheduling. This invention can avoid the irrationality and lag of manual scheduling decisions and helps to improve the production management efficiency of enterprises. Attached Figure Description

[0147] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0148] Figure 1 This is a schematic diagram of the optimization method for the distributed reentrant workshop dynamic scheduling problem provided in an embodiment of the present invention;

[0149] Figure 2 This is a schematic diagram of the optimization algorithm flow for the distributed reentrant workshop dynamic scheduling problem provided in an embodiment of the present invention;

[0150] Figure 3 This is a block diagram of an optimization device for the distributed reentrant workshop dynamic scheduling problem provided in an embodiment of the present invention;

[0151] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0152] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0153] like Figure 1As shown, this embodiment of the invention provides an optimization method for the distributed reentrant workshop dynamic scheduling problem, which can be implemented by electronic devices. Figure 1 The flowchart shown illustrates an optimization method for the distributed reentrant shop floor dynamic scheduling problem. This method's processing flow may include the following steps:

[0154] S1. Obtain the processing information of the workshop to be scheduled.

[0155] The processing information may include the number of workpieces to be processed, processing time, number of parallel machines, allocation results, and delivery time.

[0156] The processing time includes the processing time of the first section, the processing time of the second section, and the processing time of the third section; the allocation result is the result of the workpiece being allocated to different factories in the third section.

[0157] S2. Define the state feature vector, action set, and reward function for reinforcement learning.

[0158] Optionally, the set state feature vector in S2 may include six state feature vectors defined based on changes in the workpiece in the machine buffer; and two empty states defined when the machine buffer queue is less than or equal to 1.

[0159] The state feature vector may include state feature vector 1, state feature vector 2, state feature vector 3, state feature vector 4, state feature vector 5, and state feature vector 6.

[0160] Among them, the state feature vector 1 represents the ratio of the maximum remaining processing time of the workpiece in the current buffer queue to the average remaining processing time of all workpieces.

[0161] In one feasible implementation, the state feature vector 1(f) s (1) It can be represented by the following formula (1):

[0162]

[0163] Where h is the factory index, h = 1, 2, ..., f; k and c are both process indices, k = 1, 2, ..., K i (h) ; i is the index of the workpiece in buffer j, i = 1, 2, ..., BN j (h) j is the agent index, j = 1, 2, ..., m (h) ;BN j (h) Let m be the number of workpieces in the buffer of the j-th agent in factory h; (h) K represents the number of agents in factory h; i(h) This represents the total number of processes for workpiece i in factory h; Let be the processing time of workpiece i in process k in factory h.

[0164] State feature vector 2 represents the ratio of the minimum remaining processing time of the workpiece in the buffer queue to the average remaining processing time of all workpieces.

[0165] In one feasible implementation, the state feature vector 2(f) s (2) The following formula (2) can be used:

[0166]

[0167] State feature vector 3 represents the average waiting time of all artifacts in the buffer queue.

[0168] In one feasible implementation, the state feature vector 3(f) s (3) It can be represented by the following formula (3):

[0169]

[0170] in, The waiting time of workpiece i in front of the j-th intelligent agent in factory h.

[0171] State feature vector 4 represents the mean relaxation time of all workpieces in the buffer queue.

[0172] In one feasible implementation, the state feature vector 4(f) s (4) The result can be shown in equation (4) below:

[0173]

[0174] Where, d i Let i be the delivery date for workpiece i; Let i be the decision variable. If workpiece i is immediately moved to the next work section after being processed in factory h, in the dynamic distributed reentrant scheduling problem, For factory h r With the next section factory h r+1 The distance is V; the workpiece transport speed between adjacent sections is V; R is the set of sections, R = {1, 2, 3}; t is the current time.

[0175] The state feature vector 5 represents the ratio of the maximum to the average processing time of the current process in the buffer queue.

[0176] In one feasible implementation, the state feature vector 5(f) s(5) The following formula (5) can be used:

[0177]

[0178] The state feature vector 6 represents the ratio of the minimum to the average processing time of the current process in the buffer queue.

[0179] In one feasible implementation, the state feature vector 6(f) s (6) The following formula (6) can be used:

[0180]

[0181] An empty state can include empty state 1 and empty state 2.

[0182] Among them, empty state 1 indicates that the machine buffer is empty.

[0183] An empty state 2 indicates that there is one and only one workpiece in the machine buffer.

[0184] Optionally, the action set in S2 may include: minimum relaxation time rule (SLT), high response ratio priority rule (HRN), shortest remaining processing time rule (LWR), shortest processing time priority rule (SPT), longest processing time priority rule (LPT), and earliest delivery date rule (EDD).

[0185] The SLT rule means selecting the workpiece with the shortest relaxation time. If the relaxation times are the same, then one of the workpieces with the same relaxation time is randomly selected.

[0186] In one feasible implementation, the SLT rule can be represented by the following equations (7) and (8):

[0187] i * =argmin{ST i |1≤i≤BN j (h)} (7)

[0188]

[0189] The HRN rule means selecting the workpiece with the largest ratio of the sum of waiting time and processing time to the processing time. If the ratios are the same, then one of the workpieces with the same ratio is randomly selected.

[0190] In one feasible implementation, the HRN rule can be represented by the following equation (9):

[0191]

[0192] The LWR rule means selecting the workpiece with the shortest remaining processing time; if the remaining processing times are the same, then the workpiece with the earliest arrival time is selected.

[0193] In one feasible implementation, the LWR rule can be represented by the following equation (10):

[0194]

[0195] The SPT rule selects the workpiece with the shortest processing time in the current process. If the processing times in the current processes are the same, the workpiece with the earliest arrival time is selected.

[0196] In one feasible implementation, the SPT rule can be represented by the following equation (11):

[0197]

[0198] The LPT rule means selecting the workpiece with the longest processing time in the current process. If the processing times in the current processes are the same, then the workpiece with the earliest arrival time is selected.

[0199] In one feasible implementation, the LPT rule can be as shown in equation (12):

[0200]

[0201] The EDD rule means selecting the workpiece with the earliest delivery date; if the delivery dates are the same, then the workpiece with the earliest arrival time is selected.

[0202] In one feasible implementation, the EDD rule can be as shown in equation (13):

[0203] i * =argmin{d i |1≤i≤BN j (h)} (13)

[0204] Optionally, the reward function reflects the immediate reward obtained by each agent after choosing an action, and the learning objective of the agents in the system during the iteration process is to maximize the cumulative reward. The reward function can be set as shown in the following equation (14):

[0205]

[0206] Where makespan is the total processing time of the workpiece sequence.

[0207] S3. Based on the processing information, state feature vector, action set, reward function, and improved double-Q learning algorithm, the workshop scheduling scheme is obtained.

[0208] Optionally, such as Figure 2As shown, the workshop scheduling scheme obtained in S3 based on processing information, state feature vectors, action sets, reward functions, and the improved double-Q learning algorithm includes:

[0209] S301. Set the parameters of the improved double-Q learning algorithm.

[0210] The parameters may include, but are not limited to, the learning rate α, the discount factor γ, the number of clusters m, and the number of training iterations. t The iteration count is iter, the experience sharing rate is K, and the greedy strategy parameter is e.

[0211] S302, Training phase: Obtain the agent's state.

[0212] S303, Initialize each agent j Table and Initialize the common Q-value table; let k0 = 0.99, i = 1.

[0213] S304. If i < iter*0.2, then execute S305; if i ≥ iter*0.2, then execute S311.

[0214] In one feasible implementation, if i < iter * 0.2, it is the "exploration" phase of the iteration, and the double Q table learning is performed.

[0215] S305. Calculate the value of the state feature vector of each agent j based on the agent's state, cluster the values ​​of the state feature vectors of each agent j to obtain the clustering state of the agents; look up the common Q-value table based on the clustering state to obtain the common Q-value; and use the value of the common Q-value table with a probability of 1-K. Replace the current agent j The values ​​in the table and The values ​​in the table are used to obtain the values ​​of the replaced agent j. Table and surface.

[0216] S306, according to the replacement surface, The table shows the actions selected by the greedy strategy.

[0217] Optionally, in S306, according to the replacement surface, The table and the actions selected by the greedy strategy include:

[0218] Choose an action with a probability of 1-e Or randomly select action a with probability e. t The action chosen is obtained;

[0219] Among them, actions The selection process includes:

[0220] Randomly select the replacement Table or The table serves as the basis for action selection;

[0221] If you select the replacement If the table is used as the basis for action selection, then the replacement will be selected according to the greedy strategy. Maximum value of the table Action as action

[0222] If you select the replacement If the table is used as the basis for action selection, then the replacement will be selected according to the greedy strategy. Maximum value of the table Action as action

[0223] In one feasible implementation, the optimal action is selected with a probability of 1-e, and the action is randomly selected with a probability of e, as shown in equations (15) and (16) below:

[0224]

[0225]

[0226] Where the step size μ = 0.05.

[0227] Furthermore, if the action is selected As an action to be chosen, the replacement is randomly selected. Table or The table serves as the basis for action selection; if selected... If the table is used as the basis for action selection, then the replacement will be selected according to the greedy strategy. Maximum value of the table Action as action If you choose The table serves as the basis for action selection; therefore, the replacement is selected according to the greedy strategy. Maximum value of the table Action as action

[0228] S307. Execute the selected action to transition the agent's state to state s. t+1 And calculate the immediate reward r based on the reward function.

[0229] S308, Update the replacement agent j selected in S306 Table or The table is updated, and the public Q-value table is updated.

[0230] Optionally, the update in S308 is the replacement agent j selected in S306. The table includes:

[0231] When agent j is in a non-empty state, the replaced agent j is updated according to the following equations (17) and (18). surface:

[0232]

[0233]

[0234] Where t represents time s t The state at time t; a t The action at time t; Choose action a for agent j at time t. t The corresponding Q A The table shows the values; α is the learning rate; r t γ is the instant reward at time t; γ is the discount factor; s t+1 The state at time t+1; a t+1 For s t+1 Actions in a certain state; For s t+1 Q of agent j in state A The maximum value in the table; Let s be the time t+1. t+1 In the given state, agent j selects action a. t+1 The corresponding Q B The values ​​in the table.

[0235] In one feasible implementation, if in step S306 the agent j selects... If the table is used as the basis for action selection, the calculation method is the same as the selection method described above. The table serves as the same basis for action selection.

[0236] When agent j is in empty state 2, update the replaced agent j according to the following formula (19). surface:

[0237]

[0238] Optionally, the updated common Q-value table in S308 includes:

[0239] Public Q-value table and intelligent agent Table and The average value of the table is empirically shared and calculated using the following formulas (20) and (21):

[0240]

[0241]

[0242] Among them, s t+1 The state at time t+1; Indicates in s t+1 Time makes The action with the largest value; m (h) This represents the current number of intelligent agents in the factory.

[0243] S309, Let K = 0.99 i *k0; Updates the state of agent j.

[0244] S310. Determine whether all workpieces have been scheduled. If they have been scheduled, set i = i + 1 and proceed to S304. If they have not been scheduled, proceed to S305.

[0245] S311, Random selection Table replacement Table or selection Table replacement The table is obtained after assimilation.

[0246] S312. Calculate the value of the state feature vector of each agent, cluster the values ​​of the state feature vectors of each agent to obtain the cluster state of the agents; obtain the common Q value based on the cluster state, and adopt the value of the common Q value table with a probability of 1-K. replace The value of the table.

[0247] S313, Replace with The table selects the maximum value based on the greedy strategy. action Choose an action with a probability of 1-e Or randomly select action a with probability e. t , thus obtaining the action to choose.

[0248] In one feasible implementation, the optimal action is selected with a probability of 1-e, and the action is randomly selected with a probability of e, as shown in equations (22) and (23) below:

[0249]

[0250]

[0251] Where the step size μ = 1.

[0252] S314. Execute the selected action to transition the agent's state to state s. t+1 And calculate the immediate reward r based on the reward function.

[0253] S315, Updated and Replaced Table and common Q-value table; let K = 0.99 i *k0; Updates the state of the agent.

[0254] Optionally, the updated replacement in S315 The table includes:

[0255] When agent j is in a non-empty state, the replaced data is updated according to the following formula (24). surface:

[0256]

[0257] When agent j is in empty state 2, the replaced state is updated according to the following formula (25). surface:

[0258]

[0259] in, For time t The value of the table; s t The state at time t; a t Let r be the action at time t; α be the learning rate; r be the action at time t; α be the learning t γ is the instant reward at time t; γ is the discount factor; s t+1 The state at time t+1; a t+1 For s t+1 The action selected in the current state; Let s be the time t+1. t+1 In state The maximum value of the action.

[0260] Optionally, the updated common Q-value table in S315 includes:

[0261] Public Q-value table and intelligent agent Experience sharing is performed, calculated using the following formula (26):

[0262]

[0263] Among them, s t+1 The state at time t+1; Indicates in s t+1 Time makes The action with the largest value; m (h) This represents the current number of intelligent agents in the factory.

[0264] In one feasible implementation, the current factory m is used. (h) An intelligent agent in s t+1 Moment Action The corresponding maximum Q value update QC .

[0265] S316. Determine whether all workpieces have been scheduled. If they have been scheduled, proceed to S317. If they have not been scheduled, proceed to S312.

[0266] S317. If i < iter, then let i = i + 1 and proceed to execute S312; if i ≥ iter, then output the workshop scheduling scheme obtained through iteration.

[0267] In this embodiment of the invention, an optimization model is established for the multi-stage distributed reentrant shop floor dynamic scheduling problem. Considering the efficiency of processing and production, the evaluation indicators for scheduling are set as total delay and completion time. Reinforcement learning is used for real-time scheduling, transforming the distributed reentrant shop floor dynamic scheduling problem into a reinforcement learning problem. A reward function is designed based on the two evaluation indicators of completion time and total delay. The classic Q-learning algorithm cannot maintain good performance in complex production environments with multiple agents, multiple work sections, multiple factories, and reentrancy. Therefore, the classic Q-learning algorithm is improved to obtain an improved double-Q learning algorithm. To avoid Q-learning overestimating action values, a double-Q table learning mechanism is designed in the early stage of iteration; and the greedy selection coefficient is updated with a small step size in the early stage to increase the probability of exploration; a single-Q table learning mechanism is adopted in the later stage of iteration to increase the step size of the greedy selection; a multi-agent experience interaction mechanism is used throughout the entire process of the algorithm. This aims to respond promptly to dynamic events such as machine failures and random arrival of workpieces during the manufacturing process, achieving real-time scheduling. This invention can avoid the irrationality and lag of manual scheduling decisions, and helps to improve the production management efficiency of enterprises.

[0268] like Figure 3 As shown, this embodiment of the invention provides an optimization apparatus 300 for the distributed reentrant shop floor dynamic scheduling problem. This apparatus 300 is applied to an optimization method for implementing the distributed reentrant shop floor dynamic scheduling problem. The apparatus 300 includes:

[0269] The acquisition module 310 is used to acquire the processing information of the workshop to be scheduled. The processing information includes the number of workpieces to be processed, processing time, number of parallel machines, allocation result and delivery time. The processing time includes the processing time of the first section, the processing time of the second section and the processing time of the third section. The allocation result is the result of the workpieces being allocated to different factories in the third section.

[0270] The configuration module 320 is used to configure the state feature vector, action set, and reward function for reinforcement learning.

[0271] The output module 330 is used to obtain the workshop scheduling scheme based on the processing information, state feature vector, action set, reward function and improved double-Q learning algorithm.

[0272] Optionally, the set state feature vector includes the set state feature vector and the empty state.

[0273] The state feature vectors include state feature vector 1, state feature vector 2, state feature vector 3, state feature vector 4, state feature vector 5, and state feature vector 6.

[0274] Among them, the state feature vector 1 represents the ratio of the maximum remaining processing time of the workpiece in the current buffer queue to the average remaining processing time of all workpieces.

[0275] State feature vector 2 represents the ratio of the minimum remaining processing time of the workpiece in the buffer queue to the average remaining processing time of all workpieces.

[0276] State feature vector 3 represents the average waiting time of all artifacts in the buffer queue.

[0277] State feature vector 4 represents the mean relaxation time of all workpieces in the buffer queue.

[0278] The state feature vector 5 represents the ratio of the maximum to the average processing time of the current process in the buffer queue.

[0279] The state feature vector 6 represents the ratio of the minimum to the average processing time of the current process in the buffer queue.

[0280] The empty state includes empty state 1 and empty state 2.

[0281] Among them, empty state 1 indicates that the machine buffer is empty.

[0282] An empty state 2 indicates that there is one and only one workpiece in the machine buffer.

[0283] Optionally, the action set includes: minimum relaxation time rule (SLT), high response ratio priority rule (HRN), shortest remaining processing time rule (LWR), shortest processing time priority rule (SPT), longest processing time priority rule (LPT), and earliest delivery date rule (EDD).

[0284] The SLT rule means selecting the workpiece with the shortest relaxation time. If the relaxation times are the same, then one of the workpieces with the same relaxation time is randomly selected.

[0285] The HRN rule means selecting the workpiece with the largest ratio of the sum of waiting time and processing time to the processing time. If the ratios are the same, then one of the workpieces with the same ratio is randomly selected.

[0286] The LWR rule means selecting the workpiece with the shortest remaining processing time; if the remaining processing times are the same, then the workpiece with the earliest arrival time is selected.

[0287] The SPT rule selects the workpiece with the shortest processing time in the current process. If the processing times in the current processes are the same, the workpiece with the earliest arrival time is selected.

[0288] The LPT rule means selecting the workpiece with the longest processing time in the current process. If the processing times in the current processes are the same, then the workpiece with the earliest arrival time is selected.

[0289] The EDD rule means selecting the workpiece with the earliest delivery date; if the delivery dates are the same, then the workpiece with the earliest arrival time is selected.

[0290] Optionally, the output module 330 is further used for:

[0291] S301. Set the parameters of the improved double-Q learning algorithm; the parameters include the learning rate α, discount factor γ, number of clusters m, and number of training iterations. t The iteration count is iter, the experience sharing rate is K, and the greedy strategy parameter is e.

[0292] S302, Obtain the state of the agent.

[0293] S303, Initialize each agent j Table and Initialize the common Q-value table; let k0 = 0.99, i = 1.

[0294] S304. If i < iter*0.2, then execute S305; if i ≥ iter*0.2, then execute S311.

[0295] S305. Calculate the value of the state feature vector of each agent j based on the agent's state, cluster the values ​​of the state feature vectors of each agent j to obtain the clustering state of the agents; obtain the common Q-value based on the clustering state; and use the value of the common Q-value table with a probability of 1-K. Replace the current agent j The values ​​in the table and The values ​​in the table are used to obtain the values ​​of the replaced agent j. Table and surface.

[0296] S306, according to the replacement surface, The table shows the actions selected by the greedy strategy.

[0297] S307. Execute the selected action to transition the agent's state to state s. t+1 And calculate the immediate reward r based on the reward function.

[0298] S308, Update the replacement agent j selected in S306 Table or The table is updated, and the public Q-value table is updated.

[0299] S309, Let K = 0.99 i *k0; Updates the state of agent j.

[0300] S310. Determine whether all workpieces have been scheduled. If they have been scheduled, set i = i + 1 and proceed to S304. If they have not been scheduled, proceed to S305.

[0301] S311, Random selection Table replacement Table or selection Table replacement The table is obtained after assimilation.

[0302] S312. Calculate the value of the state feature vector of each agent, cluster the values ​​of the state feature vectors of each agent to obtain the cluster state of the agents; obtain the common Q value based on the cluster state, and adopt the value of the common Q value table with a probability of 1-K. replace The value of the table.

[0303] S313, Replace with The table selects the maximum value based on the greedy strategy. action Choose action a with a probability of 1-e * Or, with probability e, randomly select action a. t , thus obtaining the action to choose.

[0304] S314. Execute the selected action to transition the agent's state to state s. t+1 And calculate the immediate reward r based on the reward function.

[0305] S315, Updated and Replaced Table and common Q-value table; let K = 0.99 i *k0; Updates the state of the agent.

[0306] S316. Determine whether all workpieces have been scheduled. If they have been scheduled, proceed to S317. If they have not been scheduled, proceed to S312.

[0307] S317. If i < iter, then let i = i + 1 and proceed to execute S312; if i ≥ iter, then output the workshop scheduling scheme obtained through iteration.

[0308] Optionally, the output module 330 is further used for:

[0309] Choose an action with a probability of 1-e Or randomly select action a with probability e. t , thus obtaining the action to choose.

[0310] Among them, actions The selection process includes:

[0311] Randomly select the replacement Table or The table serves as the basis for action selection.

[0312] If you select the replacement If the table is used as the basis for action selection, then the replacement will be selected according to the greedy strategy. Maximum value of the table Action as action

[0313] If you select the replacement If the table is used as the basis for action selection, then the replacement will be selected according to the greedy strategy. Maximum value of the table Action as action

[0314] Optionally, the output module 330 is further used for:

[0315] When agent j is in a non-empty state, the replaced agent j is updated according to the following equations (1) and (2). surface:

[0316]

[0317]

[0318] Where t represents time s t The state at time t; a t The action at time t; Choose action a for agent j at time t. t The corresponding Q A The table shows the values; α is the learning rate; r t γ is the instant reward at time t; γ is the discount factor; s t+1 The state at time t+1; a t+1 For s t+1 Actions in a certain state; For s t+1 Q of agent j in state A The maximum value in the table; Let s be the time t+1. t+1 In the given state, agent j selects action a. t+1 The corresponding Q B The values ​​in the table.

[0319] When agent j is in empty state 2, the replaced agent j is updated according to the following formula (3). surface:

[0320]

[0321] Optionally, the output module 330 is further used for:

[0322] Public Q-value table and intelligent agent Table and The average value of the table is empirically shared and calculated using the following formulas (4) and (5):

[0323]

[0324]

[0325] Among them, s t+1 The state at time t+1; Indicates in s t+1 Time makes The action with the largest value; m (h) This represents the current number of intelligent agents in the factory.

[0326] Optionally, the output module 330 is further used for:

[0327] When agent j is in a non-empty state, the replaced data is updated according to the following formula (6). surface:

[0328]

[0329] When agent j is in empty state 2, the replaced state is updated according to the following formula (7). surface:

[0330]

[0331] in, For time t The value of the table; s t The state at time t; a t Let r be the action at time t; α be the learning rate; r be the action at time t; α be the learning t γ is the instant reward at time t; γ is the discount factor; s t+1 The state at time t+1; a t+1 For s t+1 The action selected in the current state; Let s be the time t+1. t+1 In state The maximum value of the action.

[0332] Optionally, the output module 330 is further used for:

[0333] Public Q-value table and intelligent agent Experience sharing is performed, calculated using the following formula (8):

[0334]

[0335] Among them, s t+1 The state at time t+1; Indicates in s t+1 Time makes The action with the largest value; m (h) This represents the current number of intelligent agents in the factory.

[0336] In this embodiment of the invention, an optimization model is established for the multi-stage distributed reentrant shop floor dynamic scheduling problem. Considering the efficiency of processing and production, the evaluation indicators for scheduling are set as total delay and completion time. Reinforcement learning is used for real-time scheduling, transforming the distributed reentrant shop floor dynamic scheduling problem into a reinforcement learning problem. A reward function is designed based on the two evaluation indicators of completion time and total delay. The classic Q-learning algorithm cannot maintain good performance in complex production environments with multiple agents, multiple work sections, multiple factories, and reentrancy. Therefore, the classic Q-learning algorithm is improved to obtain an improved double-Q learning algorithm. To avoid Q-learning overestimating action values, a double-Q table learning mechanism is designed in the early stage of iteration; and the greedy selection coefficient is updated with a small step size in the early stage to increase the probability of exploration; a single-Q table learning mechanism is adopted in the later stage of iteration to increase the step size of the greedy selection; a multi-agent experience interaction mechanism is used throughout the entire process of the algorithm. This aims to respond promptly to dynamic events such as machine failures and random arrival of workpieces during the manufacturing process, achieving real-time scheduling. This invention can avoid the irrationality and lag of manual scheduling decisions, and helps to improve the production management efficiency of enterprises.

[0337] Figure 4 This is a schematic diagram of the structure of an electronic device 400 provided in an embodiment of the present invention. The electronic device 400 can vary considerably due to differences in configuration or performance. It may include one or more central processing units (CPUs) 401 and one or more memories 402. The memory 402 stores at least one instruction, which is loaded and executed by the processor 401 to implement the following optimization method for the distributed reentrant shop floor dynamic scheduling problem:

[0338] S1. Obtain the processing information of the workshop to be scheduled; the processing information includes the number of workpieces to be processed, processing time, number of parallel machines, allocation result and delivery time; among which, the processing time includes the processing time of the first section, the processing time of the second section and the processing time of the third section; the allocation result is the result of the workpieces being allocated to different factories in the third section;

[0339] S2. Define the reinforcement learning state feature vector, action set, and reward function;

[0340] S3. Based on the processing information, state feature vector, action set, reward function, and improved double-Q learning algorithm, the workshop scheduling scheme is obtained.

[0341] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including instructions that can be executed by a processor in a terminal to complete the optimization method for the distributed reentrant workshop dynamic scheduling problem described above. For example, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0342] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0343] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for optimizing distributed re-entrant job-shop dynamic scheduling problem, characterized in that, The method comprises: S1, obtaining processing information of a to-be-scheduled workshop; the processing information comprises a to-be-processed workpiece quantity, a processing time, a parallel machine quantity, an allocation result, and a delivery time; wherein the processing time comprises a first section processing time, a second section processing time, and a third section processing time; and the allocation result is a result of allocating workpieces in the third section to different factories; S2, setting a state feature vector, an action set, and a reward function of reinforcement learning; S3, obtaining a workshop scheduling scheme according to the processing information, the state feature vector, the action set, the reward function, and an improved double Q learning algorithm; The S3 comprises: S301, set parameters of the improved double Q learning algorithm; the parameters include a learning rate , a discount factor , a number of clusters , a number of training times , a number of iterations , an experience sharing rate , and a greedy strategy parameter ; S302, obtaining an agent state; S303, initialize each agent of the table and table, initialize the common value table; let , ; S304, if S305 is executed; if S311 is executed; S305, calculating a value of a state feature vector of each agent according to the agent state, clustering the value of the state feature vector of each agent to obtain a clustered state of the agent, obtaining a common value according to the clustered state, and replacing a value in a table of the agent with a value in a common value table with a probability, to obtain a replaced table of the agent. ​​​​​​​​​​​​ S306、according to the replaced table, table and the greedy policy to get the selected action; S307, performing the selected action, transferring the agent state to a state and computing an immediate reward according to a reward function ; S308, updating the replaced agent selected in S306 of table or table, and updating the public value table; S309、Let ; update the state of the agent ; S310, judging whether all workpieces have been dispatched; if yes, then letting go to S304; if not, then going to S305. S311, randomly selecting for use table replacement table or selection for use table replacement table, resulting in a table after assimilation ; S312, calculate the value of the state feature vector of each agent at present, cluster the value of the state feature vector of each agent to obtain the cluster state of the agent; obtain the public value according to the cluster state, and replace the value of the table with the value of the public value table with the probability ​​​​​ S313, replace the... The table selects the maximum value based on the greedy strategy. action ;by The probability of selecting the action or with The probability of randomly selecting an action The action chosen is obtained; S314, performing the selected action, transferring the agent state to a state and computing an immediate reward according to a reward function ; S315、updating the replaced table and a public Q-value table; causing ; updating the state of the agent; S316, judging whether all workpieces have been scheduled, if yes, executing S317; if not, executing S312; S317, if Then let , then proceed to execute S312; if Then the iteratively obtained workshop scheduling scheme will be output.

2. The method of claim 1, wherein, The S2 comprises setting a state feature vector and an empty state; The state feature vector comprises a state feature vector 1, a state feature vector 2, a state feature vector 3, a state feature vector 4, a state feature vector 5, and a state feature vector 6; The state feature vector 1 represents a ratio of a maximum remaining processing time of a workpiece in a current buffer queue to an average remaining processing time of all workpieces; The state feature vector 2 represents a ratio of a minimum remaining processing time of a workpiece in the buffer queue to the average remaining processing time of all workpieces; The state feature vector 3 represents an average waiting time of all workpieces in the buffer queue; The state feature vector 4 represents a mean value of slack times of all workpieces in the buffer queue; The state feature vector 5 represents a ratio of a maximum value to a mean value of current process processing times in the buffer queue; The state feature vector 6 represents a ratio of a minimum value to the mean value of the current process processing times in the buffer queue; The empty state comprises an empty state 1 and an empty state 2; The empty state 1 represents that a machine buffer is empty; The empty state 2 represents that there is only one workpiece in the machine buffer.

3. The method of claim 1, wherein, The action set in the S2 comprises a slack time minimum SLT rule, a high response ratio priority HRN rule, a remaining processing time length minimum LWR rule, a shortest processing time priority SPT rule, a longest processing time priority LPT rule, and an earliest delivery deadline EDD rule; The SLT rule represents selecting a workpiece with a minimum slack time, and if the slack times are the same, randomly selecting one workpiece with the same slack time; The HRN rule represents selecting a workpiece with a maximum sum of waiting time and processing time length, or a maximum ratio of waiting time to processing time length, and if the ratios are the same, randomly selecting one workpiece with the same ratio; The LWR rule represents selecting a workpiece with a minimum remaining processing time length, and if the remaining processing times are the same, selecting a workpiece with an earliest arrival time; The SPT rule represents selecting a workpiece with a minimum current process processing time, and if the current process processing times are the same, selecting a workpiece with an earliest arrival time; The LPT rule means selecting a workpiece with the longest processing time in a current process, and if the processing time in the current process is the same, selecting a workpiece with the earliest arrival time; The EDD rule means selecting a workpiece with the earliest delivery time, and if the delivery time is the same, selecting a workpiece with the earliest arrival time.

4. The method of claim 1, wherein, The S306 according to the replacement in the table, The action selected by the table and the greedy policy includes: selecting an action with a probability of , or randomly selecting an action with a probability of , resulting in a selected action;​​ The action The selection process includes: The replaced table or table as a basis for action selection; If the replaced table is selected as the basis for action selection, the action with the maximum value of the replaced table is selected as the action according to a greedy policy. ​ If the replaced table is selected as the basis for action selection, the action with the maximum value of the replaced table is selected as the action according to a greedy policy. ​ 5. The method of claim 1, wherein, the updated agent selected in S306 in the S308 of the table includes: When the agent is in a non-empty state, update the replaced agent according to the following formulas (1), (2) Table: (1) (2) in, Indicates time, for Current state; for Actions at any given moment; for Time-based intelligent agent Select Action corresponding The values ​​of the table; The learning rate; for Instant rewards; Discount factor; for Current state; for Actions in a certain state; for intelligent agent of The maximum value in the table; for time intelligent agent in state Select Action corresponding The values ​​in the table; When the agent is in the empty state 2, the replaced agent is updated according to the following equation (3) Table: (3)。 6. The method of claim 1, wherein, The updating of the common Q value table in the S308 comprises: Public q-table with agent table with The average values of the tables are shared empirically, calculated by the following equations (4), (5): (4) (5) in, for Current state; Indicates in Time makes The action with the largest value; This represents the current number of intelligent agents in the factory.

7. The method of claim 1, wherein, The update in the S315 replaces the The table includes: When the agent is in a non-empty state, the replaced table is updated according to the following equation (6) (6) When the agent is in the empty state 2, the replaced table is updated according to the following equation (7) Table: (7) wherein, is the time instant the value of the table; is the time instant state; is the time instant action; is the learning rate; is the time instant immediate reward; is the discount factor; is the time instant state; is the action selected in the state; is the time instant in the state the maximum value of the table action.

8. The method of claim 1, wherein, The updating of the common Q value table in the S315 comprises: Public q-value table and agent Experience sharing is performed, calculated by the following formula (8): (8) wherein, is the moment state; indicates the action that makes the value of the moment so that the value of the action is the largest; is the current number of factory agents.

9. An apparatus for optimizing a distributed re-entrant job-shop dynamic scheduling problem, said apparatus being configured to implement the method for optimizing a distributed re-entrant job-shop dynamic scheduling problem according to any one of claims 1 to 8, characterized in that, The device comprises: An acquisition module is configured to acquire processing information of a to-be-scheduled workshop; the processing information comprises a to-be-processed workpiece quantity, processing time, a parallel machine quantity, an allocation result, and a delivery time; the processing time comprises a first section processing time, a second section processing time, and a third section processing time; and the allocation result is a result of allocation of a workpiece to different factories in the third section. A setting module is configured to set a state feature vector, an action set, and a reward function of reinforcement learning; An output module is configured to obtain a workshop scheduling scheme according to the processing information, the state feature vector, the action set, the reward function, and an improved double Q learning algorithm.

Citation Information

Patent Citations

  • Decision optimization method for energy storage in transaction market based on double-Q learning algorithm

    CN110598925A

  • Optimization method and device for distributed reentrant workshop scheduling problem

    CN114118699A