A Scheduling and Maintenance Optimization Method and System Based on Embedded Reinforcement Learning

By adopting the scheduling and maintenance optimization method of embedded reinforcement learning in the environment where order dynamic arrival, the scheduling and maintenance agent and feature selection agent are built to learn the optimal strategy and feature selection policy, the problem of insufficient joint optimization performance of scheduling and maintenance in the existing technology is solved, and efficient and low-cost equipment management is achieved.

CN118735200BActive Publication Date: 2025-05-27DONGHUA UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410869849.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-01
Publication Date
2025-05-27
Estimated Expiration
2044-07-01

AI Technical Summary

Technical Problem

The prior art is difficult to achieve joint optimization of scheduling and maintenance in an environment where orders are dynamically arrived, and it is difficult to accurately characterize the dynamically changing environment in the predefined state space, resulting in insufficient performance.

Method used

The scheduling and maintenance optimization method based on embedded reinforcement learning is adopted, and by constructing scheduling and maintenance agents and feature selection agents, the optimal scheduling and maintenance strategy and state feature selection policy are learned to achieve joint optimization of dynamic scheduling and maintenance.

Benefits of technology

It improves equipment reliability and reduces production costs, realizes efficient joint optimization of scheduling and maintenance, and improves the production efficiency and competitiveness of manufacturing enterprises.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118735200B_ABST
    Figure CN118735200B_ABST
Patent Text Reader

Abstract

The present invention relates to a scheduling and maintenance optimization method and system based on embedded reinforcement learning. The method comprises the following steps: collecting production operation process and machine maintenance historical data; constructing a scheduling and maintenance agent and a feature selection agent; constructing a scheduling and maintenance Markov decision process; constructing a feature selection Markov decision process; the feature selection agent interacts with the feature selection Markov decision process and learns an optimal state feature selection strategy; the scheduling and maintenance agent interacts with the scheduling and maintenance Markov decision process and learns an optimal scheduling and maintenance optimization strategy; deploying and executing the feature selection agent and the scheduling and maintenance agent to perform scheduling and maintenance optimization. The system includes a scheduling and maintenance controller, a scheduling and maintenance agent, and a feature selection agent. It solves the problem of insufficient performance caused by the difficulty in accurately characterizing the dynamic environment in the production environment where orders arrive dynamically, realizes the joint optimization of scheduling and maintenance, improves the reliability of equipment, and reduces production costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a scheduling and maintenance optimization method and system based on embedded reinforcement learning, belonging to the technologies of big data processing, artificial intelligence, and scheduling and maintenance optimization. Background Art

[0002] Production scheduling and equipment maintenance are two indispensable factors in production links and scenarios in industries and manufacturing where there are high-speed rotating, bottleneck, high-precision requirements, and high equipment reliability requirements. They jointly ensure the smooth progress of the production process and improve the production efficiency and competitiveness of enterprises. In production scheduling, if equipment maintenance is carried out frequently, while greatly improving the reliability of the equipment, it will reduce production efficiency and bring higher maintenance costs. On the contrary, low-frequency equipment maintenance will lead to the degradation of equipment reliability and even failures, thus interrupting production and resulting in additional equipment repair costs. The existing time-based maintenance strategy separates scheduling and maintenance, making it difficult to achieve integrated optimization of scheduling and maintenance. Therefore, it is of great practical significance to reasonably optimize the scheduling of dynamically arriving production orders and consider the maintenance activities of equipment to achieve the joint optimization of equipment reliability and production cost.

[0003] Currently, for the problem of joint optimization of scheduling and maintenance, maintenance activities are optimized simultaneously with production tasks as a kind of task. Therefore, the essence of the integrated optimization problem of production scheduling and flexible maintenance with dynamically arriving orders is a kind of dynamic scheduling optimization problem. The existing dynamic scheduling optimization methods mainly include proactive scheduling, pre-reactive scheduling, and online scheduling, etc. Among them, proactive scheduling quantifies uncertainties at the cost of sacrificing scheduling performance and considers them when formulating a scheduling plan to improve the robustness of the plan. However, it is difficult to accurately calculate and predict the dynamically arriving orders. Therefore, the proactive scheduling method is difficult to handle the problems of dynamic scheduling and flexible maintenance. The pre-reactive scheduling method pre-generates an initial scheduling plan and re-optimizes the affected tasks and resources after the occurrence of dynamic disturbances. However, the response time of the pre-reactive scheduling method for re-optimizing the scheduling after being disturbed increases with the increase of the problem scale, and re-generating the scheduling plan may cause a large adjustment to the system and affect the efficient production of the system. The fully reactive scheduling method performs real-time online scheduling and maintenance optimization through heuristic rules or reinforcement learning methods, can make real-time decisions based on the local information of the current production, and can better adapt to actual production. Among them, the heuristic rules have low performance due to their short-sighted characteristics, while the reinforcement learning method can train the optimal scheduling and maintenance optimization strategy and make the optimal decision-making action by comprehensively observing the current state information, which is more suitable for solving the scheduling and optimization problems.

[0004] However, the performance of reinforcement learning methods is positively correlated with the accurate representation of the environment. That is, the more accurately the state characteristics of the environment are observed, the higher the performance of the agent. Currently, the definition of the state space in reinforcement learning relies on expert experience for manual design, which easily leads to redundant or insufficient selected features. Additionally, due to the dynamic changes in the production environment along with the execution of the agent's decisions and the dynamic arrival of various types of product orders with irregular arrival times, the dynamic nature of the environment changes is increased. This results in different data distributions for task types, quantities, arrival times, etc., and the equipment working conditions and reliability also change continuously with the execution of production and maintenance activities. Therefore, it is difficult for the predefined state space to accurately represent the dynamically changing environment, thereby reducing the algorithm performance. There is an urgent need to study an adaptive feature selection method for reinforcement learning state characteristics to enable the agent to autonomously select state characteristics and improve the algorithm performance. Currently, feature selection methods mainly include three categories: filter, wrapper, and embedded. Among them, the filter method separates feature selection from algorithm execution, making it difficult to select effective key features and improve algorithm performance. The wrapper method integrates feature selection with the algorithm and selects a suitable feature subset based on the algorithm performance as feedback. However, the wrapper method requires long-term online training and is difficult to meet the high timeliness requirements of actual production. The embedded method integrates feature selection with the algorithm and trains the feature selection model offline, and in actual production, it outputs the optimal state feature subset in real time according to the changes in the environment operating state, which is more suitable for the state space design of scheduling and maintenance agents.

[0005] Therefore, in order to address issues such as dynamic scheduling considering flexible maintenance in an environment with dynamic order arrivals and the difficulty of accurately representing the dynamically changing environmental state by a predefined state space, it is necessary to study an embedded reinforcement learning method for joint optimization of dynamic scheduling and maintenance to improve equipment reliability and reduce production costs. Summary of the Invention

[0006] Aiming at the problem of insufficient performance caused by the difficulty of accurately representing the dynamic environment in the joint optimization method of production scheduling and maintenance under the condition of dynamic order arrivals, a scheduling and maintenance optimization method and system based on embedded reinforcement learning are proposed. It realizes high-reliability and low-cost scheduling and maintenance decisions in the actual production process, considers flexible equipment maintenance activities, realizes joint optimization of scheduling and maintenance, improves equipment reliability, and reduces production costs, which has important theoretical significance and practical value for improving the production efficiency and competitiveness of manufacturing enterprises.

[0007] The technical solution of the present invention is as follows:

[0008] A scheduling and maintenance optimization method based on embedded reinforcement learning, comprising the following steps: collecting historical data on production operations and machine maintenance; constructing a scheduling and maintenance agent and a feature selection agent; constructing a scheduling and maintenance Markov decision process; constructing a feature selection Markov decision process; the feature selection agent interacting with the feature selection Markov decision process and learning an optimal state feature selection strategy; the scheduling and maintenance agent interacting with the scheduling and maintenance Markov decision process and learning an optimal scheduling and maintenance optimization strategy; deploying and executing the feature selection agent and the scheduling and maintenance agent for scheduling and maintenance optimization;

[0009] Step S 1 : Collecting historical data on production operations and machine maintenance;

[0010] Among them, the historical data on production operations includes, but is not limited to, production orders, scheduling plans, historical production operation data, types of production tasks, batch sizes, arrival times, workstations assigned tasks, production sequences, start processing times, end processing times, etc.;

[0011] The historical data on machine maintenance includes, but is not limited to, historical operation data such as equipment service life, production batches, production time, maintenance time, maintenance time intervals, equipment operating status, etc.;

[0012] Step S 2 : Constructing a scheduling and maintenance agent and a feature selection agent;

[0013] Among them, the scheduling and maintenance agent includes a policy network and a target network, which are constructed by a deep neural network; the policy network is used to select the most appropriate scheduling or maintenance activity, and the target network is used to update the parameters of the policy network;

[0014] The feature selection agent includes an actor network and a critic network; the actor network and the critic network are constructed by a deep neural network; the actor network is used to select key features as the input of the scheduling and maintenance agent, and the critic network is used to update the parameters of the actor network;

[0015] Step S 3 : Constructing a scheduling and maintenance Markov decision process;

[0016] Among them, the scheduling and maintenance Markov decision process is to transform the scheduling and maintenance optimization problem into a sequential decision-making problem, that is, a Markov decision process, mainly including the definition of the state space, action space, and reward function of scheduling and maintenance;

[0017] Define the state space of scheduling and maintenance: It is the key feature of the scheduling and maintenance Markov decision process output after triggering the feature selection agent to monitor the scheduling and maintenance Markov decision process;

[0018] Define the action space of scheduling and maintenance: It includes two actions of selecting the current task to be processed or machine maintenance. The scheduling and maintenance action space is as follows:

[0019] A 2 = [1, 2,..., n, n + 1]

[0020] Among them, A 2 represents the scheduling and maintenance action space, [1, 2,..., n] represents the current task to be processed, n represents the number of tasks to be processed, and n + 1 represents machine maintenance;

[0021] Define the scheduling and maintenance reward function: It evaluates any action selected by the scheduling and maintenance agent from the scheduling and maintenance action space A 2 . The scheduling and maintenance reward function is composed of switching cost, maintenance cost, repair cost, and the reliability of the machine. The total reward of the scheduling and maintenance agent is as follows:

[0022]

[0023] In the formula, R 2 represents the total reward of the scheduling and maintenance agent, m is the number of parallel machines, T is the total cycle of scheduling and maintenance, r jt represents the reliability of machine j at time t, C csc , C ssc , C rc , C mc respectively represent the production interruption cost, product switching cost for different specifications, machine repair cost, and machine maintenance cost in the same batch;

[0024] Among them, the reliability of all machines is fitted through the Weibull distribution. The calculation formula of the reliability r jt is as follows:

[0025]

[0026] In the formula, T as is the age of the machine, β is the shape parameter of the Weibull distribution, and η is the scale parameter of the Weibull distribution;

[0027] The production interruption cost, product switching cost for different specifications, equipment repair cost, equipment maintenance cost, etc. are calculated through the production interruption time, material loss per unit time, and unit material cost. The calculation formulas are as follows in sequence;

[0028] Ccsc = T cst * U i * C i

[0029] C ssc = T sst * U i * C i

[0030] C rc = T rt * U i * C i

[0031] C mc = Tm t * U i * C i

[0032] wherein, T cst , T sst , T rt , T mt successively represent the preparation time, product changeover time, equipment repair time, and equipment maintenance time between two subtasks in the same batch, U j represents the unit time output of product i, and C j represents the unit material cost of product i;

[0033] Since the total reward R 2 is calculated from all the immediate rewards r 2,t of the scheduling and maintenance agent in the period T, as shown in the following formula:

[0034] R 2 = r 2,1 + y * r 2,2 + γ 2 * r 2,3 +... + γ t-1 * r 2,t +...

[0035] wherein, r 2,1 , r 2,2 , r 2,3 , … successively represent the immediate rewards of the immediate reward r 2,t at times t = 1, t = 2, t = 3, …, γ is the discount reward, and when γ = 7, R 2 = r 2,1 + r 2,2 + r 2,3 +... + r 2,t +...;

[0036] Therefore, the immediate rewards of the scheduling and maintenance agent are divided into the following situations;

[0037] At decision-making moment \(t\), the scheduling and maintenance agent selects a maintenance action, which will increase the cost while improving the reliability. The reward at this time is shown as follows:

[0038]

[0039] At decision-making moment \(t\) 2 the scheduling and maintenance agent chooses to continue producing the current batch and the machine does not break down during the subsequent production process, only increasing the setup cost and reducing the reliability. The reward at this time is shown as follows:

[0040]

[0041] At decision-making moment \(t\) 3 the scheduling and maintenance agent chooses to continue producing the current batch and the machine breaks down during the subsequent production process, which will increase the setup cost and repair cost and reduce the reliability. The reward at this time is shown as follows:

[0042]

[0043] At decision-making moment \(t\) 4 the scheduling and maintenance agent chooses to change the batch, which will cause a conversion cost, but will replace the components of the machine to restore the reliability to its original state. The reward at this time is shown as follows:

[0044]

[0045] Step S 4 : Construct a feature selection Markov decision process;

[0046] The feature selection Markov decision process transforms the key feature selection problem of the scheduling and maintenance Markov decision process into a sequential decision problem, that is, the Markov decision process, which mainly includes the definition of the feature selection state space, action space and reward function;

[0047] Define the feature selection state space, which contains all production operation processes and machine maintenance data;

[0048] Define the feature selection action space, including two actions: selection and elimination of all production operation processes and machine maintenance data;

[0049] Define the feature selection reward function, which is obtained by taking the key features contained in the selected production operation process and machine maintenance data as the input of the scheduling and maintenance agent and training the scheduling and maintenance agent, and using the optimized scheduling and maintenance performance of the scheduling and maintenance agent as the total reward of the feature selection agent;

[0050] R1 = R 2,t

[0051] wherein, R 1 represents the total reward of the feature selection agent, and R 2,t represents the total reward of the scheduling and maintenance agent trained with the key features output by the feature selection agent at time t;

[0052] Step S 5 : The feature selection agent interacts with the feature selection Markov decision process and learns the optimal state feature selection strategy;

[0053] The interaction means that the feature selection agent observes the current production operation process and machine maintenance data s 1,t , selects and executes an action a 1,t , observes the production operation process and machine maintenance data s changed after executing the action 1,t+1 , and obtains the key feature training scheduling and maintenance agent after selecting and executing all actions to obtain the total reward R 1 ;

[0054] The optimal feature selection strategy is that the feature selection agent interacts with the feature selection Markov decision process and updates the parameters of the deep neural network in the feature selection agent through the Actor-Critic algorithm guided by self-imitation learning;

[0055] Self-imitation learning is to use the interaction data contained in the better trajectories obtained by the feature selection agent exploring the solution space as expert experience, and sample by combining all the explored trajectory data to update the parameters of the feature selection agent to solve the sparse reward problem faced by the feature selection agent and accelerate the learning and convergence speed of the feature selection agent;

[0056] The Actor-Critic algorithm is a reinforcement learning algorithm that combines policy gradients and value networks, used to train the feature selection agent to learn the optimal feature selection strategy, that is, to train the deep neural network in the feature selection agent to obtain the optimal network parameters;

[0057] The parameter update formula of the feature selection agent of the Actor-Critic algorithm guided by self-imitation learning is as follows in sequence:

[0058]

[0059] In the formula, θ i , θ i+7 are the parameters of the actor network of the feature selection agent before and after update respectively, represents the objective function J of the actor network Q (θ i) gradient, η Q is the learning rate of the actor network, where,

[0060]

[0061] In the formula, J Q (θ) represents the objective function of the actor network when the actor network parameters are θ, Q θ (s 1,t , a 1,t ) represents the predicted value when the actor network selects action a 1,t at the time when the parameters are θ and the state is s 1,t , y is the target value, respectively represent the expected values of the mean square error calculated using the interaction data included in the better and worse trajectories;

[0062]

[0063] In the formula, r is the reward value when the feature selection agent actor network selects action a 1,t at the time when the parameters are θ and the state is s 1,t , α is the learning rate, represents the action a 1,t+1 with the highest probability selected according to its own policy by the feature selection agent actor network when the state is s 1,t+1 ;

[0064]

[0065] In the formula, are the parameters of the feature selection agent critic network before and after update respectively, represents the objective function of the critic network gradient, η Π is the learning rate of the critic network, where,

[0066]

[0067] In the formula, represents the objective function of the critic network when the critic network parameters are θ, represents the action a 1,t with the highest probability selected according to its own policy by the feature selection agent critic network when the state is s 1,t ;

[0068] Step S 6: The scheduling and maintenance agent interacts with the scheduling and maintenance Markov decision process and learns the optimal scheduling and maintenance optimization strategy;

[0069] The interaction is that the scheduling and maintenance agent observes the current production operation process and machine maintenance data to obtain the state s by selecting the key features output by the agent according to the features 2,t , selects and executes the action a through its own policy combined with the adaptive action selection mechanism 2t , obtains the immediate reward r feedback from the scheduling and maintenance Markov decision process 2,t , observes the state data s transferred after executing the action 2,t+7 ;

[0070] The optimal scheduling and maintenance strategy is the neural network parameters trained and updated by the scheduling and maintenance agent and the scheduling and maintenance Markov decision process to accumulate experience data through the Double DQN algorithm;

[0071] The adaptive action selection mechanism is to solve the problem of non - selectable actions when the scheduling and maintenance agent selects and executes actions, so as to improve the global exploration ability of the agent for the solution space. It mainly determines the set of selectable actions by judging the current selectable actions to generate an action mask matrix, and then the scheduling and maintenance agent selects the selectable actions according to the action mask matrix combined with the ε - greedy strategy;

[0072] The ε - greedy strategy is shown as the following formula:

[0073]

[0074] In the formula, random represents randomly selecting a selectable action, represents selecting the action with the largest Q value, R(0, 1) represents randomly taking a value between 0 and 1, ε is the greedy coefficient, and θ is the parameter of the policy network in the scheduling and maintenance agent;

[0075] The Double DQN algorithm is a value - based reinforcement learning algorithm used to train the scheduling and maintenance agent to learn the optimal scheduling and maintenance strategy. The error Loss calculation formula of the scheduling and maintenance agent is as follows:

[0076] Loss = / / y i - Q(s 2,t , a 2,t ; ω) / / 2

[0077]

[0078] In the formula, ω is the parameter of the policy network in the scheduling and maintenance agent, To schedule and maintain the parameters of the evaluation network in the agent, Q(S 2,t , a 2,t ; ω) is the Q value of the agent choosing action a 2,t according to its own policy parameter ω in state s 2,t . The error of parameter ω is calculated by gradient descent, and the calculation formula is as follows:

[0079]

[0080] In the formula, represents the gradient of the error Loss of the policy network in the scheduling and maintenance agent under parameter ω;

[0081] Step S 7 : Deploy and execute the feature selection agent and the scheduling and maintenance agent for scheduling and maintenance optimization;

[0082] Deploy the trained feature selection agent and scheduling and maintenance agent to the production workshop, and automatically trigger the feature selection agent to select the key state features of the dynamic environment and use the scheduling and maintenance agent to output the optimal scheduling or maintenance activities in the current state through real-time monitoring of the production and machine operation process data.

[0083] Preferably, in step S 2 , the scheduling and maintenance agent includes a policy network and a target network, and the policy network and the target network can be constructed by a fully connected neural network.

[0084] Preferably, in step S 2 , the feature selection agent includes an actor network and a critic network; the actor network and the critic network are constructed by a fully connected neural network.

[0085] A scheduling and maintenance optimization system based on embedded reinforcement learning is used for the scheduling and maintenance optimization method based on embedded reinforcement learning. The system includes a scheduling and maintenance controller, a scheduling and maintenance agent, and a feature selection agent;

[0086] The scheduling and maintenance controller, by real-time monitoring of the scheduling and maintenance process, accepts the trigger signal transmitted by the production operation environment and transmits the signal to the scheduling and maintenance agent to trigger the generation of the scheduling and maintenance plan;

[0087] The scheduling and maintenance agent, after receiving the trigger signal of the state data transmission from the scheduling and maintenance controller, sends a signal to the feature selection agent, accepts the real-time selection of the current key state features by the feature selection agent, and uses itself to learn the optimal scheduling and maintenance strategy to select and execute the optimal scheduling and maintenance actions;

[0088] After receiving the trigger signal transmitted by the scheduling and maintenance agent, the feature selection agent detects the state features of the scheduling and maintenance environment and outputs the key feature subset in the current environment state by using the learned optimal feature selection strategy to transmit to the scheduling and maintenance agent.

[0089] The beneficial effects of the present invention are as follows:

[0090] 1. Aiming at the scheduling and flexible maintenance problems of dynamically arriving orders, the present invention proposes an optimization method for scheduling and flexible maintenance based on the embedded Double DQN (Double Deep Q-Network) algorithm. By exploring the complex coupling relationship between equipment reliability and repair cost, conversion cost, and maintenance cost through agents, the optimal scheduling and maintenance strategy is learned to achieve the joint optimization of scheduling and maintenance, thereby effectively reducing production costs and improving equipment reliability.

[0091] 2. Aiming at the problem that the pre-defined state space in the scheduling and maintenance agent is difficult to accurately represent the dynamic environment and reduces the algorithm performance, the present invention provides a state feature selection method based on the improved Actor-Critic algorithm. By learning the optimal feature selection strategy, the autonomous selection of key state features of the scheduling and maintenance environment is realized to further improve the performance of the scheduling and maintenance method.

[0092] 3. Aiming at the problem that the scheduling and maintenance agent has insufficient exploration ability for the complex scheduling and maintenance solution space, the present invention provides an adaptive action selection mechanism to improve the global exploration ability of the scheduling and maintenance agent for the solution space and improve the algorithm performance.

[0093] 4. Aiming at the sparse reward problem faced in the learning process of the state feature selection strategy, the present invention provides a self-imitation learning algorithm to guide the feature selection agent to accelerate the learning performance and convergence speed by imitating the excellent experience trajectories explored by itself. Brief Description of the Drawings

[0094] Figure 1 It is a schematic diagram of the implementation steps of an optimization method for scheduling and maintenance based on embedded reinforcement learning provided by the present invention;

[0095] Figure 2 It is a schematic diagram of the logic of an optimization method for scheduling and maintenance based on embedded reinforcement learning provided by the present invention;

[0096] Figure 3 It is a schematic diagram of the system architecture of an optimization method for scheduling and maintenance based on embedded reinforcement learning provided by the present invention. Detailed Embodiment

[0097] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. These embodiments are implemented on the premise of the technical solution of the present invention, and provide detailed implementation manners and specific operation processes. However, the protection scope of the present invention is not limited to the following embodiments.

[0098] A scheduling and maintenance optimization method based on embedded reinforcement learning includes the following steps: collecting production operation process and machine maintenance historical data; constructing a scheduling and maintenance agent and a feature selection agent; constructing a scheduling and maintenance Markov decision process; constructing a feature selection Markov decision process; the feature selection agent interacts with the feature selection Markov decision process and learns the optimal state feature selection strategy; the scheduling and maintenance agent interacts with the scheduling and maintenance Markov decision process and learns the optimal scheduling and maintenance optimization strategy; deploying and executing the feature selection agent and the scheduling and maintenance agent for scheduling and maintenance optimization.

[0099] Specifically, the embodiments of the present invention will be further described below with reference to the accompanying drawings.

[0100] In one embodiment, as shown in Figure 1 In the embodiment of the present invention, a scheduling and maintenance optimization method based on embedded reinforcement learning is provided. The logic of the method is shown in Figure 2 , and the method includes:

[0101] Step S 1 : Collect production operation process and machine maintenance historical data;

[0102] Among them, the production operation process historical data includes, but is not limited to, production orders, scheduling plans, production operation historical data, production task types, batches, arrival times, workstations assigned tasks, production sequences, start processing times, end processing times, etc.

[0103] The machine maintenance historical data includes, but is not limited to, historical operation data such as equipment service life, production batches, production times, maintenance times, maintenance time intervals, and equipment operation status.

[0104] Step S 2 : Construct a scheduling and maintenance agent and a feature selection agent;

[0105] Among them, the scheduling and maintenance agent includes a policy network and a target network. The policy network and the target network are constructed by a deep neural network, such as a fully connected neural network. The policy network is used to select the most appropriate scheduling or maintenance activity, and the target network is used to update the parameters of the policy network.

[0106] The feature selection agent includes an actor network (actor network) and a critic network (critic network). The actor network and the critic network are constructed by a deep neural network, such as a fully connected neural network. The actor network is used to select key features as the input of the scheduling and maintenance agent, and the critic network is used to update the parameters of the actor network.

[0107] Step S 3 : Construct a scheduling and maintenance Markov decision process (Markov decision process);

[0108] Among them, the scheduling and maintenance Markov decision process is to transform the scheduling and maintenance optimization problem into a sequential decision problem, that is, the Markov decision process, which mainly includes the definition of the state space, action space, and reward function of scheduling and maintenance.

[0109] Define the state space of scheduling and maintenance: It is the key feature of the scheduling and maintenance Markov decision process output after triggering the feature selection agent to monitor the scheduling and maintenance Markov decision process.

[0110] Define the action space of scheduling and maintenance: It includes two actions of selecting the current task to be processed or machine maintenance. The scheduling and maintenance action space is as follows:

[0111] A 2 =[1, 2,..., n, n + 1]

[0112] Among them, A 2 represents the scheduling and maintenance action space, [1, 2,..., n] represents the current tasks to be processed, n represents the number of tasks to be processed, and n + 1 represents machine maintenance.

[0113] Define the scheduling and maintenance reward function: It is to evaluate any action selected by the scheduling and maintenance agent from the scheduling and maintenance action space A 2 . The scheduling and maintenance reward function is composed of switching cost, maintenance cost, repair cost, and the reliability of the machine. The total reward of the scheduling and maintenance agent is as follows:

[0114]

[0115] In the formula, R 2 represents the total reward of the scheduling and maintenance agent, m is the number of parallel machines, T is the total cycle of scheduling and maintenance, r jt represents the reliability of machine j at time t, C csc 、C ssc 、C rc 、C mcrespectively represent the production interruption cost, product switching cost for different specifications, machine repair cost, and machine maintenance cost in the same batch.

[0116] Among them, the reliability of all machines is fitted by the Weibull distribution, and the reliability r jt has the following calculation formula:

[0117]

[0118] In the formula, T as is the service age of the machine, β is the shape parameter of the Weibull distribution, and η is the scale parameter of the Weibull distribution.

[0119] The production interruption cost, product switching cost for different specifications, equipment repair cost, equipment maintenance cost, etc. are calculated through the production interruption time, material loss per unit time, and unit material cost. The calculation formulas are as follows in sequence.

[0120] C csc = T cst * U i * C i

[0121] C ssc = T sst * U i * C i

[0122] C rc = T rt * U i * C i

[0123] C mc = T mt * U i * C i

[0124] In the formula, T cst 、T sst 、T rt 、T mt respectively represent the preparation time, product switching time, equipment repair time, and equipment maintenance time between two subtasks in the same batch, U j represents the output per unit time of product i, and C j represents the unit material cost of product i.

[0125] Since the total reward R 2 is calculated from all the immediate rewards r 2,t obtained by the scheduling and maintenance agent in the period T, as shown in the following formula:

[0126] R 2 = r 2,1+γ*r 2,2 +γ 2 *r 2,3 +...+γ t-1 *r 2,t +...

[0127] where r 2,1 、r 2,2 、r 2,3 、… represent the immediate rewards r 2,t at times t = 1, t = 2, t = 3, … respectively. γ is the discounted reward. When γ = 1, R 2 = r 2,1 + r 2,2 + r 2,3 +...+ r 2,t +...

[0128] Therefore, the immediate rewards of the scheduling and maintenance agent are divided into the following situations.

[0129] At decision time t, when the scheduling and maintenance agent selects a maintenance action, it will increase the cost while improving the reliability. The reward is as shown in the following formula:

[0130]

[0131] At decision time t 2 when the scheduling and maintenance agent selects to continue producing the current batch and the machine does not fail during the subsequent production process, only the setup cost is increased and the reliability is reduced. The reward is as shown in the following formula:

[0132]

[0133] At decision time t 3 when the scheduling and maintenance agent selects to continue producing the current batch and the machine fails during the subsequent production process, the setup cost and the repair cost are increased and the reliability is reduced. The reward is as shown in the following formula:

[0134]

[0135] At decision time t 4 when the scheduling and maintenance agent selects to change batches, it will incur a switching cost, but it will replace the components of the machine to restore the reliability to its original state. The reward is as shown in the following formula:

[0136]

[0137] Step S 4 : Construct a feature selection Markov decision process;

[0138] The feature selection Markov decision process transforms the key feature selection problem of the scheduling and maintenance Markov decision process into a sequential decision problem, that is, the Markov decision process, which mainly includes the definition of the feature selection state space, action space, and reward function.

[0139] Define the feature selection state space, which contains all production operation processes and machine maintenance data.

[0140] Define the feature selection action space, including two actions: selecting and eliminating all production operation processes and machine maintenance data.

[0141] Define the feature selection reward function. After using the key features contained in the selected production operation processes and machine maintenance data as the input for the scheduling and maintenance agent and training the scheduling and maintenance agent, the optimized scheduling and maintenance performance of the scheduling and maintenance agent is used as the total reward for the feature selection agent.

[0142] R 1 =R 2 . t

[0143] Where, R 1 represents the total reward of the feature selection agent, and R 2,t represents the total reward of the scheduling and maintenance agent trained with the key features output by the feature selection agent at time t.

[0144] Step S 5 : The feature selection agent interacts with the feature selection Markov decision process and learns the optimal state feature selection strategy;

[0145] The interaction is that the feature selection agent observes the current production operation process and machine maintenance data s 1,t , selects and executes an action a 1,t , observes the production operation process and machine maintenance data s 1,t+1 changed after executing the action, and trains the scheduling and maintenance agent with the key features obtained after selecting and executing all actions to obtain the total reward R 1 .

[0146] The optimal feature selection strategy is that the feature selection agent interacts with the feature selection Markov decision process and updates the parameters of the deep neural network in the feature selection agent through the Actor-Critic algorithm guided by self-imitation learning.

[0147] Self-imitation learning is to use the interaction data contained in the better trajectories obtained by the feature selection agent exploring the solution space as expert experience, and sample by combining all the explored trajectory data to update the parameters of the feature selection agent to solve the sparse reward problem faced by the feature selection agent and accelerate the learning and convergence speed of the feature selection agent.

[0148] The Actor-Critic algorithm is a reinforcement learning algorithm that combines policy gradients and value networks, used to train the feature selection agent to learn the optimal feature selection strategy, that is, to train the deep neural network in the feature selection agent to obtain the optimal network parameters.

[0149] The parameter update formula of the feature selection agent of the Actor-Critic algorithm guided by self-imitation learning is as follows:

[0150]

[0151] In the formula, θ i , θ i+7 are the parameters of the actor network of the feature selection agent before and after the update, respectively. represents the gradient of the objective function J Q (θ i ), η Q is the learning rate of the actor network. Among them,

[0152]

[0153] In the formula, J Q (θ) represents the objective function of the actor network when the actor network parameters are θ, Q θ (s 1,t , a 1,t ) represents the predicted value when the actor network selects the action a 1,t when the parameters are θ and the state is s 1,t , y is the target value, respectively represent the expected values of the mean square errors calculated using the interaction data contained in the better and worse trajectories.

[0154]

[0155] In the formula, r is the reward value when the actor network of the feature selection agent selects the action a 1,t when the parameters are θ and the state is s 1,t , α is the learning rate, represents the action a 1,t+1 with the highest probability selected according to its own policy by the actor network of the feature selection agent when the state is s 1,t+1 .

[0156]

[0157] Wherein, are the parameters of the critic network of the feature selection agent before and after update respectively, represents the objective function of the critic network the gradient of η Π is the learning rate of the critic network, where,

[0158]

[0159] Wherein, represents the objective function of the critic network when the parameters of the critic network are θ, represents that the critic network of the feature selection agent selects the action a with the highest probability according to its own policy 1,t when the state is s 1,t .

[0160] Step S 6 : The scheduling and maintenance agent interacts with the scheduling and maintenance Markov decision process and learns the optimal scheduling and maintenance optimization strategy;

[0161] The interaction is that the scheduling and maintenance agent observes the state s by observing the current production operation process and machine maintenance data according to the key features output by the feature selection agent 2,t , selects and executes the action a through its own policy combined with the adaptive action selection mechanism 2,t , obtains the immediate reward r feedback by the scheduling and maintenance Markov decision process 2,t , observes the state data s transferred after executing the action 2,t+7 .

[0162] The optimal scheduling and maintenance strategy is the neural network parameters trained and updated by the scheduling and maintenance agent and the scheduling and maintenance Markov decision process through interacting to accumulate experience data by the Double DQN algorithm.

[0163] The adaptive action selection mechanism is to solve the problem of non - selectable actions when the scheduling and maintenance agent selects and executes actions, so as to improve the global exploration ability of the agent for the solution space. It mainly determines the set of selectable actions to generate an action mask matrix by judging the current selectable actions, and then the scheduling and maintenance agent selects the selectable actions according to the action mask matrix combined with the ε - greedy strategy.

[0164] The ε - greedy strategy is shown in the following formula:

[0165] ​

[0166] In the formula, random represents randomly selecting an optional action, represents selecting the action with the largest Q value, R(0, 1) represents randomly taking a value between 0 and 1, ε is the greedy coefficient, and θ is the parameter of the policy network in the scheduling and maintenance agent.

[0167] The Double DQN algorithm is a value-based reinforcement learning algorithm used to train the scheduling and maintenance agent to learn the optimal scheduling and maintenance strategy. The error Loss calculation formula of the scheduling and maintenance agent is as follows:

[0168] Loss = / / y i -Q(s 2,t , a 2,t ; ω) / / 2

[0169]

[0170] In the formula, ω is the parameter of the policy network in the scheduling and maintenance agent, is the parameter of the evaluation network in the scheduling and maintenance agent, Q(S 2,t , a 2,t ; ω) is the Q value of the agent selecting action a 2,t under state s according to its own policy parameter ω, and the error of parameter ω is calculated through gradient descent. The calculation formula is as follows: 2,t

[0171]

[0172] In the formula, represents the gradient of the error Loss of the policy network in the scheduling and maintenance agent under parameter ω.

[0173] Step S 7 : Deploy and execute the feature selection agent and the scheduling and maintenance agent for scheduling and maintenance optimization.

[0174] Deploy the trained feature selection agent and scheduling and maintenance agent to the production workshop, and automatically trigger the feature selection agent to select the key state features of the dynamic environment and use the scheduling and maintenance agent to output the optimal scheduling or maintenance activity in the current state by real-time monitoring of the production and machine operation process data.

[0175] In the second aspect, the present invention provides a dynamic scheduling and maintenance optimization system based on reinforcement learning. Refer to Figure 3 , the system includes a scheduling and maintenance controller, a scheduling and maintenance agent, and a feature selection agent.

[0176] The scheduling and maintenance controller, by monitoring the scheduling and maintenance process in real time, receives the trigger signal transmitted by the production operation environment and transmits the signal to the scheduling and maintenance agent to trigger the generation of the scheduling and maintenance plan.

[0177] The scheduling and maintenance agent, after receiving the trigger signal transmitted by the scheduling and maintenance controller's status data, sends a signal to the feature selection agent, receives the real-time selection of the current key state features by the feature selection agent, and uses itself to learn the optimal scheduling and maintenance strategy to select and execute the optimal scheduling and maintenance actions.

[0178] After receiving the trigger signal transmitted by the scheduling and maintenance agent, the feature selection agent detects the state features of the scheduling and maintenance environment and outputs the key feature subset in the current environment state to the scheduling and maintenance agent by using the learned optimal feature selection strategy.

[0179] The above-described embodiments merely represent one implementation manner of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several variations and improvements can be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.

Claims

1. A scheduling and maintenance optimization method based on embedded reinforcement learning, characterized in that: The following steps are involved: Collect production operation process and machine maintenance history data; build scheduling and maintenance agents and feature selection agents; build scheduling and maintenance Markov decision-making processes; Construct a feature selection Markov decision process; the feature selection agent interacts with the feature selection Markov decision process and learns the optimal state feature selection strategy; the scheduling and maintenance agent interacts with the scheduling and maintenance Markov decision process and learns the optimal scheduling and maintenance optimization strategy; deploy and execute the feature selection agent and the scheduling and maintenance agent to perform scheduling and maintenance optimization; the details are as follows: Step S1: Collecting production operation process and machine maintenance history data; Among them, historical data of the production operation process, including but not limited to production orders, scheduling plans, historical data of production operation, types of production tasks, batches, arrival times, workstations to which tasks are assigned, production sequences, start processing time, and end processing time; Machine maintenance history data, including but not limited to equipment age, production batches, production time, maintenance time, maintenance interval, and equipment operating status historical operation data; Step S2: construct a scheduling and maintenance agent and a feature selection agent; The scheduling and maintenance agent includes a policy network and a target network, which are constructed by deep neural networks. The policy network is used to select the most appropriate scheduling or maintenance activities, and the target network is used to update the parameters of the policy network. Feature selection agent, including actor network and critic network; actor network and critic network are constructed through deep neural network; actor network is used to select key features as input of scheduling and maintenance agent, and critic network is used to update parameters of actor network; Step S3: construct the scheduling and maintenance Markov decision process; Among them, the scheduling and maintenance Markov decision process is to transform the scheduling and maintenance optimization problem into a sequential decision problem, namely the Markov decision process, including the definition of the state space, action space and reward function of scheduling and maintenance; Define the state space of scheduling and maintenance: It is the key features of the scheduling and maintenance Markov decision process that are output by the triggering feature selection agent after monitoring the scheduling and maintenance Markov decision process; Define the action space for scheduling and maintenance: It includes two actions: selecting the current task to be processed or machine maintenance. The action space for scheduling and maintenance is as follows: A2=[1,2,…,n,n+1] Among them, A2 represents the scheduling and maintenance action space, [1,2,…,n] represents the current tasks to be processed, n represents the number of tasks to be processed, and n+1 represents machine maintenance; Definition of the scheduling and maintenance reward function: It evaluates any action selected by the scheduling and maintenance agent from the scheduling and maintenance action space A2. The scheduling and maintenance reward function is composed of switching cost, maintenance cost, repair cost and machine reliability. The total reward of the scheduling and maintenance agent is as follows: In the formula, R2 represents the total reward of the scheduling and maintenance agent, m is the number of parallel machines, T is the total cycle of scheduling and maintenance, and r jt represents the reliability of machine j at time t, C csc , C ssc , C rc , C mc They represent the production interruption cost in the same batch, the switching cost of products with different specifications, the machine repair cost, and the machine maintenance cost respectively; Among them, the reliability of all machines is fitted by Weibull distribution, and the reliability r jt The calculation formula is as follows: Where, T as is the service life of the machine, β is the shape parameter of the Weibull distribution, and η is the size parameter of the Weibull distribution; The production interruption cost, the switching cost of products of different specifications, the equipment repair cost, and the equipment maintenance cost are calculated by the production interruption time, the material loss per unit time, and the unit material cost. The calculation formulas are as follows: C csc =T cst *U i *C i C ssc =T sst *U i *C i C rc =T rt *U i *C i C mc =T mt *U i *C i Where, T cst , T sst , T rt , T mt represents the preparation time, product switching time, equipment repair time, and equipment maintenance time between two subtasks in the same batch, respectively. i represents the output per unit time of product i, C i represents the unit material cost of product i; Since the total reward R2 is the instantaneous reward r of the scheduling and maintenance agent in period T 2,t The calculation is as shown below: R2=r 2,1 +γ*r 2,2 +g 2 *r 2,3 +…+c t-1 *r 2,t +… Among them, r 2,1 、r 2,2 、r 2,3 , ... respectively represent the immediate reward r 2,t The instant reward at time t = 1, t = 2, t = 3, ..., γ is the discounted reward, when γ = 1, R2 = r 2,1 +r 2,2 +r 2,3 +…+r 2,t +…; Therefore, the immediate rewards of the scheduling and maintenance agents are divided into the following cases; At decision time t1, the scheduling and maintenance agent chooses the maintenance action, which increases the cost while improving reliability. The reward is as follows: At decision time t2, the scheduling and maintenance agent chooses to continue to produce the current batch and the machine does not fail in the subsequent production process, which only increases the preparation cost and reduces reliability. The reward at this time is as follows: At decision time t3, the scheduling and maintenance agent chooses to continue to produce the current batch and the machine fails in the subsequent production process, which will increase the preparation cost and maintenance cost and reduce reliability. The reward at this time is as follows: At decision time t4, the scheduling and maintenance agent chooses to change the batch, which will cause conversion costs, but will replace the components of the machine to restore the reliability to the original state. At this time, the reward is as follows: Step S4: constructing a feature selection Markov decision process; Define the feature selection state space, including all production operation process and machine maintenance data; Define the feature selection action space, including the selection and elimination of all production operation process and machine maintenance data; The feature selection reward function is defined by taking the key features contained in the selected production operation process and machine maintenance data as the input of the scheduling and maintenance agent and training the scheduling and maintenance agent, and obtaining the scheduling and maintenance optimization performance of the scheduling and maintenance agent as the total reward of the feature selection agent; R1=R 2,t Among them, R1 represents the total reward of the feature selection agent, R 2,t represents the total reward of the scheduling and maintenance agent trained with the key features output by the feature selection agent at time t; Step S5: The feature selection agent interacts with the feature selection Markov decision process and learns the optimal state feature selection strategy; Interaction is the observation of the current production process and machine maintenance data by the feature selection agent 1,t , select and perform action a 1,t , observe the changes in production operation process and machine maintenance data after executing the action 1,t+1 , select and execute the key features obtained after all actions to train the scheduling and maintenance agent to obtain the total reward R1; The optimal feature selection strategy is that the feature selection agent interacts with the feature selection Markov decision process and updates the parameters of the deep neural network in the feature selection agent through the Actor-Critic algorithm guided by self-imitation learning; The parameter update formulas of the feature selection agent of the Actor-Critic algorithm guided by self-imitation learning are as follows: In the formula, θ i ,θ i+1 Select the parameters of the actor network for the features before and after the update, represents the objective function J of the actor network Q (θ i ), η Q is the learning rate of the actor network, where In the formula, J Q (θ) represents the objective function of the actor network when the actor network parameter is θ, Q θ (s 1,t ,a 1,t ) represents the actor network with parameters θ and state s 1,t When selecting action a 1,t The predicted value at time y is the target value, They represent the expected values ​​of the mean square error calculated using the interaction data contained in the better and worse trajectories respectively; Where r is the feature selection agent actor network with parameters θ and state s 1,t When selecting action a 1,t The reward value at time , α is the learning rate, Indicates that the feature selection agent actor network is in state s 1,t+1 According to your own strategy The action with the highest probability is chosen 1,t+1 ; In the formula, The parameters of the feature selection agent critic network before and after the update, Represents the objective function of the critic network The gradient of π is the learning rate of the critic network, where In the formula, represents the objective function of the critic network when the critic network parameter is θ, Indicates that the feature selection agent critic network is in state s 1,t According to your own strategy The action with the highest probability is chosen 1,t ; Step S6: The scheduling and maintenance agent interacts with the scheduling and maintenance Markov decision process and learns the optimal scheduling and maintenance optimization strategy; The interaction is that the scheduling and maintenance agents observe the current production process and machine maintenance data according to the key features of the feature selection agent output to obtain the state s 2,t , select and execute action a through its own strategy combined with the adaptive action selection mechanism 2,t , obtain the immediate reward r fed back by the scheduling and maintenance Markov decision process 2,t , observe the state data s transferred after executing the action 2,t+1 ; The optimal scheduling and maintenance strategy is the interaction between the scheduling and maintenance agent and the scheduling and maintenance Markov decision process to accumulate experience data and train and update the neural network parameters through the Double DQN algorithm; The optional action set is determined by judging the current optional action to generate an action mask matrix. Then the scheduling and maintenance agent selects the optional action according to the action mask matrix combined with the ε-greedy strategy. The ε-greedy strategy is as follows: In the formula, random means randomly selecting an optional action. represents the selection of the action with the largest Q value, R(0,1) represents a random value between 0 and 1, ε is the greedy coefficient, and θ is the parameter of the policy network in the scheduling and maintenance agent; The error loss calculation formula of the scheduling and maintenance agent is as follows: Loss=||y i -Q(s 2,t ,a 2,t ;ω)|| 2 Where ω is the parameter of the policy network in the scheduling and maintenance agent, is the parameter of the evaluation network in the scheduling and maintenance agent, Q(s 2,t ,a 2,t ; ω) is the agent in state s 2,t Next, select action a according to its own strategy parameter ω 2,t The Q value of the parameter ω is calculated by gradient descent, and the calculation formula is as follows: In the formula, Represents the gradient of the error Loss of the policy network in the scheduling and maintenance agent under the parameter ω; Step S7: deploy and execute the feature selection agent and the scheduling and maintenance agent to perform scheduling and maintenance optimization; The trained feature selection agent and scheduling and maintenance agent are deployed in the production workshop. Through real-time monitoring of production and machine operation process data, the feature selection agent is automatically triggered to select key state features of the dynamic environment, and the scheduling and maintenance agent is used to output the optimal scheduling or maintenance activities under the current state.

2. The scheduling and maintenance optimization method based on embedded reinforcement learning according to claim 1 is characterized in that: In step S2, the scheduling and maintenance agent includes a policy network and a target network, and the policy network and the target network can be constructed by a fully connected neural network.

3. The scheduling and maintenance optimization method based on embedded reinforcement learning according to claim 1 is characterized in that: In step S2, the feature selection agent includes an actor network and a critic network; the actor network and the critic network are constructed through a fully connected neural network.

4. A scheduling and maintenance optimization system based on embedded reinforcement learning, characterized in that: A scheduling and maintenance optimization method based on embedded reinforcement learning as described in any one of claims 1 to 3, wherein the system comprises a scheduling and maintenance controller, a scheduling and maintenance agent, and a feature selection agent; The scheduling and maintenance controller monitors the scheduling and maintenance process in real time, receives the trigger signal transmitted by the production operation environment, and transmits the signal to the scheduling and maintenance agent to trigger the generation of scheduling and maintenance plans; The scheduling and maintenance agent, after receiving the trigger signal of the state data transmission from the scheduling and maintenance controller, sends a signal to the feature selection agent, accepts the real-time selection of the current key state features by the feature selection agent, and uses itself to learn the optimal scheduling and maintenance strategy to select and execute the optimal scheduling and maintenance action; After receiving the trigger signal transmitted by the scheduling and maintenance agent, the feature selection agent detects the state characteristics of the scheduling and maintenance environment, and uses the learned optimal feature selection strategy to output the key feature subset under the current environment state to pass it to the scheduling and maintenance agent.

Citation Information

Patent Citations

  • Production and maintenance coupling task allocation method and system considering equipment operation state

    CN115619171A

  • Complete vehicle manufacturing stamping resource scheduling method based on deep reinforcement learning

    CN117557016A