Production scheduling and machine maintenance real-time decision-making method based on deep reinforcement learning

Through a method based on deep reinforcement learning, a dynamic scheduling and predictive maintenance model is constructed, and the production scheduling and machine maintenance problems are decomposed, real-time adaptive optimization of the production system is achieved, and the problems of inapplicable decision-making solutions and insufficient real-time performance in the existing technology are solved, and production efficiency and economic benefits are improved.

CN120471368APending Publication Date: 2025-08-12TONGJI UNIV

Patent Information

Application Number
CN202510562257.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing technology has not been effectively combined in production scheduling and machine maintenance, resulting in the decision-making plan not being applicable to actual production scenarios, insufficient machine reliability assumptions, affecting the real-time and flexibility of decision-making, and the integrated solution of production scheduling and machine maintenance affects flexibility.

Method used

Using a method based on deep reinforcement learning, an integrated optimization model of dynamic scheduling and predictive maintenance is constructed, and the remaining service life of the machine is predicted using deep learning, which is decomposed into independent optimization of the scheduling agent and maintenance agent to generate real-time decisions.

Benefits of technology

Real-time adaptive optimization of production scheduling and machine maintenance is realized, reducing modeling bias, improving real-time and flexibility of decision-making, and reducing production delays and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471368A_ABST
    Figure CN120471368A_ABST
Patent Text Reader

Abstract

The invention relates to a production scheduling and machine maintenance real-time decision-making method based on deep reinforcement learning, and the method comprises the steps: obtaining the operation characteristics of a production workshop and the actual maintenance condition information of a machine, and constructing an integrated optimization model of dynamic scheduling and predictive maintenance based on the operation characteristics and the actual maintenance condition information; acquiring real-time state data of machine operation, predicting the residual service life of the machine by using a deep learning method based on the real-time state data, setting a first threshold value and a second threshold value of the residual service life, and dividing machine maintenance states based on the first threshold value and the second threshold value; constructing a DDQN-based intelligent agent, solving the integrated optimization model based on the intelligent agent, and generating a production scheduling and machine maintenance real-time decision; the intelligent agents comprise a scheduling intelligent agent and a maintenance intelligent agent, and the states of the intelligent agents comprise operation characteristics and remaining service life. Compared with the prior art, the method not only meets the production performance requirement of reducing the total delay, but also optimizes the maintenance cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a joint optimization technology for production and maintenance, and in particular to a production scheduling and equipment maintenance decision-making method based on deep reinforcement learning. Background Art

[0002] In modern manufacturing, production scheduling is a critical link in ensuring efficient production processes, playing a vital role in improving resource utilization and reducing production costs. However, in actual operation, equipment performance degradation is inevitable due to aging, corrosion, and wear and tear from long-term use. This reality contrasts sharply with the assumption in some studies that machines are always operational. If maintenance is not scheduled promptly, machine failures can occur, severely delaying production tasks. Furthermore, the operating status of machines not only directly affects production efficiency but also the profitability of the company. Therefore, in actual production processes, it is particularly necessary to integrate production scheduling with maintenance activities for comprehensive consideration. Therefore, solving the problem of coordinating production scheduling and machine maintenance has become a critical issue that needs to be urgently addressed in current production workshops. Previous approaches to production scheduling and machine maintenance problems often involved maintaining machines at fixed intervals in a static production environment based on a fixed production scheduling scheme. This approach often resulted in over- or under-maintenance, and because its scheduling scheme was fixed, it ignored dynamic disturbances in the production process. To address these issues, existing technologies have proposed solutions to production scheduling and machine maintenance problems by transforming a static production environment into a dynamically generated one. For example, Chinese patent application CN118735200A proposes a scheduling and flexible maintenance optimization method based on an embedded Double DQN (Dual Deep Q Network) algorithm for scheduling and flexible maintenance of dynamically arriving orders. This method addresses the over- or under-maintenance issues caused by ignoring the dynamic disturbances of random workpieces. However, it suffers from the following issues:

[0003] 1) When solving problems, the scheduling and maintenance decision-making process is directly established without dynamic association with the actual scenario, which may result in the generated decision-making solution not being applicable to the actual production scenario;

[0004] 2) The reliability of the machine in the solution is obtained by modeling based on historical data using the Weibull distribution assumption. The obtained results are not real-time enough, which affects the effectiveness of the decision-making solution.

[0005] 3) The integrated solution of production scheduling and machine maintenance is integrated into the same intelligent agent, which affects the flexibility of the solution process.

[0006] Therefore, a technical problem that needs to be solved is to provide a method for distributed production scheduling and machine maintenance decision-making solution that fits actual application scenarios and can be updated in real time. Summary of the Invention

[0007] The purpose of the present invention is to overcome the defects of the above-mentioned existing technologies and provide a real-time decision-making method for production scheduling and machine maintenance based on deep reinforcement learning. This method can adaptively optimize scheduling and maintenance strategies according to changes in the remaining service life during machine degradation, provide support for real-time decision-making in dynamic environments, and improve the economic benefits and operational efficiency of the production system.

[0008] The purpose of the present invention can be achieved by the following technical solutions:

[0009] The present invention provides a method for real-time decision-making for production scheduling and machine maintenance based on deep reinforcement learning. The method performs the following steps on a production workshop where workpieces dynamically arrive to generate real-time decisions for production scheduling and machine maintenance in the production workshop, including:

[0010] Obtaining the operating characteristics of the production workshop and the actual maintenance status information of the machines, and constructing an integrated optimization model for dynamic scheduling and predictive maintenance based on the operating characteristics and actual maintenance status information; the integrated optimization model includes a cost optimization function and constraints;

[0011] Acquiring real-time status data of the machine operation, predicting the remaining useful life of the machine using a deep learning method based on the real-time status data, setting a first threshold and a second threshold for the remaining useful life, and classifying the machine maintenance status based on the first threshold and the second threshold;

[0012] A DDQN-based intelligent agent is constructed, and the integrated optimization model is solved based on the intelligent agent to generate real-time decisions for production scheduling and machine maintenance; the intelligent agent includes a scheduling intelligent agent and a maintenance intelligent agent, and the state of the intelligent agent includes operating characteristics and remaining service life.

[0013] As a preferred technical solution, the cost optimization function expression is:

[0014] min:f=CY T +CY M ,

[0015] Among them, CT T represents the total delay cost of processing the workpiece, and T k represents the delay time of the workpiece, l represents the number of workpieces to be processed, T k =max(0,(C k -D k )) represents workpiece J k The delay time, C k Indicates the workpiece completion time, D k Indicates the delivery time of the workpiece; CTM represents the total maintenance cost of the machine, and CT M =CT CM +CT IM , represents the corrective maintenance cost of the machine, n represents the total number of machines in the production workshop, h i Indicates machine M i Corrective maintenance times, Indicates machine M i Corrective maintenance costs; represents the imperfect maintenance cost of the machine, u i Indicates machine M i The number of non-perfect maintenance, w i Indicates the number of machines M per unit time i The cost of imperfect maintenance, Indicates non-perfect maintenance time.

[0016] As a preferred technical solution, the calculation method of the non-perfect maintenance time is:

[0017]

[0018] Among them, b i represents the linear coefficient, a i Indicates machine M i Basic maintenance time, T x Indicates the running time at which the remaining service life of the machine reaches the first threshold, T t Indicates the current time.

[0019] As a preferred technical solution, the constraints include:

[0020] The completion time of each process of the workpiece needs to be earlier than the start time, the expression is: C kmi >S kmi , where C kmi Indicates workpiece J k The mth process O km On machine M i Completion time, S kmi Indicates process O km On machine M i Start time on

[0021] The workpiece is processed according to the process sequence, and the expression is: S kmi >C k(m-1)i′ , C k(m-1)i′ Indicates workpiece J k The m-1th process O k(m-1) On machine M i′ Completion time on

[0022] The maintenance of each machine is only carried out before or after the workpiece is processed. The expression is: (MS i >C kmi )&(SE i k′mi ), MS i Indicates machine M i Maintenance start time, ME i Indicates machine M i Maintenance end time; C kmi Indicates workpiece ME k The mth process O km On machine M i Completion time on k′mi Indicates workpiece J k′ The mth process O k′m On machine M i Start time on

[0023] The processing of each step is a continuous process, and the expression is: C kmi =S kmi +p kmi , p kmi Indicates process O km On machine M i Processing time on

[0024] Each process can only be performed on one machine, and the expression is: X kmi Indicates process O km Assigned to machine M i The status parameter of process O is 1. km Assigned to machine M i , otherwise its value is 0;

[0025] Each machine only processes one process in the same time period, and the expression is: d k Indicates workpiece J k The number of processes; l represents the number of workpieces to be processed;

[0026] Each time the machine performs only one task, which includes maintenance and workpiece processing. The maintenance includes imperfect maintenance and corrective maintenance. The expression is: ∑ i∈n (X kmi +Q i +Z i )≤1,X kmi Indicates process O km Assigned to machine M i State parameter; Q i Indicates the state where the machine performs non-perfect maintenance. When Q​i =1 indicates that imperfect maintenance is performed, otherwise Q i =0; Z i Indicates the state of the machine performing corrective maintenance. i =1 indicates that imperfect maintenance is performed, otherwise Z i =0;

[0027] Only when the workpiece reaches the corresponding machine, the process is executed, and the expression is: (C k1i -p k1i -A k )X k1i ≥0, A k Indicates the time when the workpiece arrives at the corresponding machine; p k1i Indicates the processing time of the first step of the workpiece; C k1i Indicates the completion time of the first process of the workpiece;

[0028] Each machine cannot handle multiple different types of processes at the same time. The expression is: (S kmi ,C kmi )∩(S k′m′i ,C k′m′i )=φ,S k′m′i and C k′m′i Represents workpiece J k′ The m′th process O k′m′ Processing start and end time; S kmi and C kmi Represents workpiece J k The mth process O km The processing start and end time.

[0029] As an optimal technical solution, it is characterized in that the state space of the scheduling agent is constructed based on the operating characteristics, including the workpiece processing queue operation state and the process operation state; its action space includes: the shortest workpiece process time, the shortest time required to complete the remaining process, the minimum progress ratio of the workpiece processing process and the earliest delivery time; its reward is calculated based on the delay time of the workpiece and the time in the process queue state.

[0030] As a preferred technical solution, the workpiece processing queue operation status includes: the average difference and minimum difference between the current time when the workpiece is in the queue operation and the delivery time, the total time, average time, minimum time and maximum time required for the remaining processes when the workpiece is in the queue operation, the total time, average time, minimum time and maximum time required for the next process when the workpiece is in the queue operation, the ratio of the maximum time when the workpiece is in the queue operation to the average time when the workpiece is in the process processing, the ratio of the minimum time when the workpiece is in the queue operation to the average time when the workpiece is in the process processing, and the sum of the maximum slack time, minimum slack time, average slack time and slack time when the workpiece is in the queue operation;

[0031] The process operation status includes: the number of workpieces being processed, the number of workpieces in queue, the average completion rate of the process processing and the lateness rate of the process processing.

[0032] As a preferred technical solution, the reward calculation expression of the scheduling agent is:

[0033]

[0034] Among them, R km represents the reward of the scheduling agent; T k Indicates the delay time of the workpiece; AT k Indicates the average delay time of each process; Q km Indicates that the workpiece enters the mth process O km Waiting time; r k Indicates workpiece J k The number of processes.

[0035] As a preferred technical solution, the state space of the maintenance agent includes the quantity characteristics, time characteristics and remaining service life of the machine maintenance; its action space is the state of whether the machine is maintained, including no maintenance required, imperfect maintenance and corrective maintenance; its reward is calculated based on the maintenance cost per unit normal operating time between the current maintenance and the last maintenance.

[0036] As a preferred technical solution, the quantity characteristics include: the number of corrective maintenance, the number of imperfect maintenance between two consecutive corrective maintenance;

[0037] The time characteristics include: the interval time between the current time and the previous maintenance, the first difference time between the current remaining service life of the machine and the first threshold, the second difference time between the current remaining service life of the machine and the second threshold, and the difference between the current time and the first difference time.

[0038] As a preferred technical solution, the reward calculation method for the maintenance agent is:

[0039]

[0040] Among them, r u represents the reward for maintaining the agent; a u represents the action of maintaining the u-th decision point of the agent; DN means no maintenance action is required; CT u represents the maintenance cost at each decision point between two adjacent corrective maintenance tasks; s Indicates the time difference between two adjacent non-DN maintenance times; IT u Represents the time interval between two consecutive maintenance decision points.

[0041] Compared with the prior art, the present invention has the following beneficial effects:

[0042] 1) The present invention constructs an integrated optimization model of dynamic production scheduling and predictive maintenance based on the actual characteristics of the production workshop. The delay cost caused by workpiece delays in the production process and the maintenance cost caused by machine maintenance are used as the cost optimization function in the integrated optimization model. The present invention also comprehensively considers the workpiece processing sequence, workpiece processing requirements, and the correspondence between workpieces and production machines in the actual production process to construct constraints that conform to the actual production logic, reduce modeling deviations, and ensure that the real-time decision-making solution is more in line with the actual application scenario.

[0043] 2) Since the remaining service life of a machine is a key indicator parameter for measuring machine production, the present invention uses the remaining service life of a machine as a state feature of the machine. The present invention uses deep learning to make real-time predictions of the remaining service life based on real-time state monitoring data. Compared with the prior art method of hypothetical modeling based on historical data, the present invention can make dynamic real-time adjustments according to the state of the machine, and the adjustment is more flexible and accurate, and the state of the intelligent body is updated synchronously in real time, thereby realizing dynamic updates of decisions.

[0044] 3) The present invention functionally decomposes the problems of production scheduling and equipment maintenance, and constructs a scheduling agent and a maintenance agent for the above two problems respectively. The two can be independently expanded and optimized and make distributed decisions. Compared with the existing technology that integrates production scheduling and equipment maintenance, the method provided by the present invention is more flexible. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 is a flow chart of the method of the present invention;

[0046] Figure 2 The scheduling and maintenance Gantt chart obtained by running 4000 time units in Example 2 of the present invention;

[0047] Figure 3This is a trend chart of the remaining service life of machine number 1 after running for 4000 time units in Example 2 of the present invention. DETAILED DESCRIPTION

[0048] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0049] Unless otherwise defined, the technical or scientific terms used in this application should have the ordinary meaning understood by a person of ordinary skill in the technical field to which this application belongs. The words "one", "a", "the" and the like used in this application do not indicate a limit on quantity and may indicate the singular or plural. The terms "include", "comprise", "have" and any variations thereof used in this application are intended to cover non-exclusive inclusions; for example, a process, method, system, product or device that includes a series of steps or modules (units) is not limited to the listed steps or units, but may also include steps or units that are not listed, or may also include other steps or units that are inherent to these processes, methods, products or devices. The words "connect", "connected", "coupled" and the like used in this application are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The word "multiple" used in this application refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the objects before and after are in an "or" relationship. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific order for the objects.

[0050] Example 1

[0051] In order to solve the problems existing in the prior art, this embodiment provides a real-time decision-making method for production scheduling and machine maintenance based on deep reinforcement learning. The method performs the following steps on a production workshop where workpieces arrive dynamically to generate real-time decisions on production scheduling and machine maintenance for the production workshop. The process is as follows: Figure 1 Shown, including:

[0052] S1. Obtain the operating characteristics of the production workshop and the actual maintenance status information of the machines, and build an integrated optimization model for dynamic scheduling and predictive maintenance based on the operating characteristics and actual maintenance status information.

[0053] In this embodiment, the integrated optimization model includes a cost optimization function and constraints, wherein the cost optimization function includes the total delay cost and the total maintenance cost, wherein the cause of the delay cost is the workpiece J k The completion time is later than the delivery date.

[0054] In detail, the expression of the cost optimization function is:

[0055] min:f=CT T +CT M ,

[0056] Among them, CT T represents the total delay cost of processing the workpiece, and T k represents the delay time of the workpiece, l represents the number of workpieces to be processed, T k =max(0,(C k -D k )) represents workpiece J k The delay time, C k Indicates the workpiece completion time, D k Indicates the delivery time of the workpiece; CT M represents the total maintenance cost of the machine, and CT M =CT CM +CT IM , represents the corrective maintenance cost of the machine, n represents the total number of machines in the production workshop, h i Indicates machine M i Corrective maintenance times, Indicates machine M i Corrective maintenance costs; represents the imperfect maintenance cost of the machine, u i Indicates machine M i The number of non-perfect maintenance, w i Indicates the number of machines M per unit time i The cost of imperfect maintenance, Indicates non-perfect maintenance time.

[0057] In this embodiment, based on the limitations of the actual production process, the following nine constraints are constructed, including:

[0058] 1) The completion time of each process of the workpiece needs to be earlier than the start time, the expression is: C kmi >S kmi , where C kmi Indicates workpiece J k The mth process O km On machine M i Completion time, S kmi Indicates process O kmOn machine M i Start time on

[0059] 2) The workpiece is processed according to the process sequence, and the expression is: S kmi >C k(m-1)i′ , C k(m-1)i′ Indicates workpiece J k The m-1th process O k(m-1) On machine M i′ Completion time on

[0060] 3) The maintenance of each machine is only carried out before or after the workpiece is processed. The expression is: (MS i >C kmi )&(SE i k′mi ), MS i Indicates machine M i Maintenance start time, ME i Indicates machine M i Maintenance end time; C kmi Indicates workpiece J k The mth process O km On machine M i Completion time on k′mi Indicates workpiece J k′ The mth process O k′m On machine M i Start time on

[0061] 4) Each process is a continuous process, and the expression is: C kmi =S kmi +p kmi , p kmi Indicates process O km On machine M i Processing time on

[0062] 5) Each process can only be performed on one machine, and the expression is: X kmi Indicates process O km Assigned to machine M i The status parameter of process O is 1. km Assigned to machine M i , otherwise its value is 0;

[0063] 6) Each machine only processes one process in the same time period. The expression is: d k Indicates workpiece J k The number of processes; l represents the number of workpieces to be processed;

[0064] ​7) Each time the machine performs only one task, which includes maintenance and workpiece processing. Maintenance includes imperfect maintenance and corrective maintenance. The expression is: ∑ i∈n (X kmi +Q i +Z i )≤1,X kmi Indicates process O km Assigned to machine M i State parameter; Q i Indicates the state where the machine performs non-perfect maintenance. When Q i =1 indicates that imperfect maintenance is performed, otherwise Q i =0; Z i Indicates the state of the machine performing corrective maintenance. i =1 indicates that imperfect maintenance is performed, otherwise Z i =0;

[0065] 8) The process is executed only when the workpiece reaches the corresponding machine. The expression is: (C k1i -p k1i -A k )X k1i ≥0, A k Indicates the time when the workpiece arrives at the corresponding machine; p k1i Indicates the processing time of the first step of the workpiece; C k1i Indicates the completion time of the first process of the workpiece;

[0066] 9) Each machine cannot process multiple different types of processes at the same time. The expression is: (S kmi ,C kmi )∩(S k′m′i ,C k′m′i )=φ,S k′m′i and C k′m′i Represents workpiece J k′ The m′th process O k′m′ Processing start and end time; S kmi and C kmi Represents workpiece J k The mth process O km The processing start and end time.

[0067] S2. Obtain real-time status data of the machine operation and use deep learning methods to predict the remaining service life of the machine based on the real-time status data.

[0068] RUL (remaining useful life) is a key indicator to measure the ability of a machine to perform production tasks. As production continues, the performance of the machine will inevitably degrade. As an indicator of the degree of degradation of the machine, the remaining useful life can take adaptive maintenance measures according to the current state of the machine. In this embodiment, the types of maintenance measures are set to include no maintenance (DN), imperfect maintenance (IM) and corrective maintenance (CM). Among them, DN means that no maintenance is required and the machine is operating normally; IM assumes that the machine can be restored to a state between "new machine state" and "aged state"; CM involves replacement due to severe degradation of the machine, usually with a new machine. For ease of comparison and analysis, the remaining useful life can be normalized to the range of [0,1], which represents the ratio of the remaining useful life of the machine to the total life. Taking into account the degree of degradation of the machine, the degradation level is divided into three intervals. Each interval needs to be in a different maintenance and repair state, and the remaining useful life thresholds of interval two and interval three are set as the first threshold H respectively. x and the second threshold H y , the specific description is as follows:

[0069] Interval 1: RUL∈[H x ,1],In this interval, the machine operates normally and no maintenance activities are required.

[0070] Interval 2: RUL∈[H y ,H x ] After the machine completes the processing task, if the RUL of the machine is within the range, maintenance measures including DN, IM, and CM are adaptively selected according to the situation.

[0071] Interval 3: RUL∈[0,H y ], after the machine completes the processing task, if the machine's RUL is less than H y , the risk of machine unreliability increases significantly, so corrective maintenance CM activities must be carried out immediately.

[0072] In addition, the degree of degradation directly affects the duration of maintenance activities. Specifically, the more serious the degree of machine degradation, the longer the required non-perfect maintenance (IM) time. After the machine completes the processing task, if IM activities are selected, the maintenance time is modeled as a linear function of the time difference. In contrast, if corrective maintenance (CM) activities are chosen, the maintenance time is fixed. Therefore, the maintenance time for IM and CM can be calculated as follows.

[0073]

[0074] Among them, b i represents the linear coefficient, a i Indicates machine M iBasic maintenance time, T x Indicates the running time when the remaining service life of the machine reaches the first threshold, T t Indicates the current time. and They are machine M i Maintenance time for performing imperfect maintenance (IM) and corrective maintenance (CM) activities, and Is a fixed value.

[0075] The deep learning method used in this embodiment is GRU, and the collected historical status data is used to train the GRU model. The acquired real-time status data is processed based on the trained model to obtain the remaining service life prediction result.

[0076] In the prediction process, the real-time state data collected is first divided into multiple continuous time periods using the sliding window technology. Each time window consists of a fixed number of past observations. These observations x t ={x t1 ,…,x tn} as the input of the prediction model. The provided GRU includes an input layer, a GRU layer, and a fully connected layer. The input layer receives and processes the observation time series containing multiple variables, and then inputs it into the GRU layer. The GRU layer extracts and processes the features to capture the health status characteristics of the machine. The output of the GRU layer serves as the input of the fully connected layer, and finally outputs the remaining service life prediction result. The parameter update formula during the training process of this model is:

[0077]

[0078] Among them, σ represents the sigmoid activation function, W represents the weight matrix, b represents the bias, * represents element-by-element multiplication, and h t-1 represents the previous hidden state, represents the candidate hidden state, h t represents the final hidden state, x t Represents the current input, r t Represents the reset gate, z t represents the update gate.

[0079] S3. Build an intelligent agent based on DDQN, solve the integrated optimization model based on the intelligent agent, and obtain real-time decisions on production scheduling and machine maintenance.

[0080] This example decomposes the integrated problem of production scheduling and machine maintenance, constructing a scheduling agent and a maintenance agent respectively, and defining the corresponding state space, action space, state transition process, and reward mechanism for the two agents. The details are as follows:

[0081] (a) Scheduling Agent:

[0082] State space: The state space of the scheduling agent is constructed based on the operation characteristics, including the workpiece processing queue operation state and the process operation state. Among them, the workpiece processing queue operation state includes: the average and minimum difference between the current time of the workpiece in the queue operation and the delivery time, the total time, average time, minimum time and maximum time required for the remaining processes when the workpiece is in the queue operation, the total time, average time, minimum time and maximum time required for the next process when the workpiece is in the queue operation, the ratio of the maximum time of the workpiece in the queue operation to the average time of the workpiece in the process processing, the ratio of the minimum time of the workpiece in the queue operation to the average time of the workpiece in the process processing, and the maximum slack time, minimum slack time, average slack time and the sum of the slack time when the workpiece is in the queue operation; the process operation state includes: the number of workpieces currently in process processing, the number of workpieces in the queue operation, the average completion rate of process processing and the lateness rate of process processing.

[0083] Action space: includes four scheduling rules: shortest process time (SPT), shortest time required to complete the remaining process (LWKR), minimum progress ratio (CR) of the workpiece processing process, and earliest delivery time (EDD). Taking the shortest process time (SPT) rule as an example, it means selecting the workpiece with the shortest processing time in the queue.

[0084] State transition process: For the scheduling agent, when the machine becomes idle, it observes the current state At this decision point, a scheduling rule is selected To select a job from the machine queue for machine processing; after completing the operation of the selected job, observe the next state award Calculated after all operations of a job are completed, the reward is assigned to each action that processed an operation of that job.

[0085] Reward mechanism: For the scheduling agent, minimizing the delay cost is equivalent to minimizing the delay of the workpiece. The longer the workpiece waits for processing on the machine, the higher the probability of delay. k The delay can only be determined after the workpiece processing is completed, so at each decision point (in J k For example, the reward can be based on the delay T of the workpiece k and its waiting time Q k There are two cases for calculation: one is workpiece J k There is no delay, T k =0, Second: Workpiece J k There is a delay, T k ≠0, its expression is:

[0086]

[0087] Among them, R km represents the reward of the scheduling agent; T k Indicates the delay time of the workpiece; AT k Indicates the average delay time of each process, AT k =T k / r k , r k Indicates the number of processes; Q km Indicates that the workpiece enters the mth process O km Waiting time; r k Indicates workpiece J k The number of processes.

[0088] (b) Maintenance Agent:

[0089] State space: includes quantity characteristics, time characteristics, and remaining service life. Among them, the quantity characteristics include: the number of corrective maintenance, the number of imperfect maintenance between two consecutive corrective maintenance; the time characteristics include: the interval time between the current time and the previous maintenance, the first difference time between the current remaining service life of the machine and the first threshold, the second difference time between the current remaining service life of the machine and the second threshold, and the difference between the current time and the first difference time.

[0090] Action space: The state of whether the machine is to be maintained, including no maintenance required, imperfect maintenance and corrective maintenance.

[0091] State transition: For maintenance agents, when a machine completes the operation and its RUL value is less than the first threshold H x When the state is observed At this point, take maintenance action for the machine After the maintenance action is completed, the next state is observed The reward is calculated after the next CM activity is executed. The reward is assigned to each maintenance action between two consecutive CM activities.

[0092] Reward mechanism: For maintenance agents, maintenance costs are directly affected by the type of maintenance action performed. Corrective maintenance (CM) incurs significantly higher costs than IM maintenance. Since DN does not incur maintenance costs, it can be assigned a relatively higher reward value than the other two actions. For corrective maintenance (CM) and IM maintenance, the reward can be determined by the maintenance cost per unit of uptime between the current maintenance and the last maintenance, including three cases: first, the current action is no maintenance required; second, the previous action is no maintenance required, and the current action is other than no maintenance required; third, other cases, the expression is:

[0093]

[0094] Among them, r u represents the reward for maintaining the agent; a u represents the action of maintaining the u-th decision point of the agent; DN means no maintenance action is required; CT u represents the maintenance cost at each decision point between two adjacent corrective maintenance tasks; s Indicates the time difference between two adjacent non-DN maintenance times, IT s =t u -t l , t u Indicates the current maintenance start time, t l Indicates the end time of the previous maintenance; IT u Represents the time interval between two consecutive maintenance decision points.

[0095] After the construction of the intelligent agent is completed, it adopts centralized training and decentralized execution mode, and parameter sharing technology is used to train the scheduling and maintenance intelligent agent during the training process. The intelligent agent training is carried out in a randomly generated production environment. The workpieces arrive at the workshop randomly, and the time interval of arrival obeys the exponential distribution T~Exp(30) to enhance the versatility of the method.

[0096] Example 2

[0097] In order to verify that the method provided in the above embodiment is feasible, it is verified in this embodiment.

[0098] In detail, according to domain expert knowledge, the remaining service life of each machine after imperfect maintenance is modeled as 80% of the life after the last maintenance. After corrective maintenance, the life of the machine will be fully restored after replacement. In this embodiment, the historical maintenance parameters of 8 production machines are collected as shown in Table 1.

[0099] Table 1 Machine historical maintenance parameters

[0100]

[0101] The historical maintenance data corresponding to the above eight machines are derived from the degradation process of actual industrial production and are used to simulate machine degradation based on filter clogging when separating solid particles from gas. In this data, the flow rate and pressure difference of the machine filter are recorded throughout the service life, and the corresponding machine operating life cycles are 626, 1217, 1267, 1993, 2631, 1285, 2936 and 2819 respectively.

[0102] Based on the historical maintenance data of the eight machines, we divided the data into an 8:2 ratio, with 80% of the data used as the training set and 20% of the data used as the validation set. The GRU model parameters obtained through training are shown in Table 2. The trained GRU is used to predict the current remaining useful life based on the currently collected machine condition monitoring data.

[0103] Table 2 GRU model parameters

[0104] Batch size Sliding window size Learning rate Optimizer Number of training sessions Number of GRUs 1024 30 0.001 Adam 250 40

[0105] According to the contents of Example 1, a scheduling agent and a maintenance agent are constructed, and the predicted remaining service life is used as one of the states of the maintenance agent. In actual production scenarios, the action networks of the trained scheduling agent and maintenance agent can adaptively select scheduling and maintenance actions at each decision point, respectively, thereby realizing adaptive real-time scheduling and maintenance decisions.

[0106] In this embodiment, the simpy discrete event simulation dynamic production environment is used to train the constructed intelligent agent, and the training time is set to 100,000 time units. The DDQN parameters corresponding to the intelligent agent are shown in Table 3. During the training process, the first threshold H x and the second threshold H y are set to 0.8 and 0.2 respectively, and the processing time of the workpiece on each machine follows a uniform distribution p kmi ∈[5,40].

[0107] Table 3 DDQN parameter table

[0108]

[0109] After the training is completed, the four scheduling rules and four maintenance rules in the agent action space are combined to generate combination rules, and the four optimal scheduling-maintenance combination rules are selected (as shown in Table 4). The MT4 maintenance rule is to set the first threshold H x Set to 0.15, the second threshold H y Set to 0.1, if the RUL of the machine is lower than 0.15, then randomly select an action from DN (no maintenance required) and IM (imperfect maintenance). If the RUL of the machine is lower than 0.1, then immediately perform CM (corrective maintenance). Compare with the method in Example 1 and run for 4000 time units. The results are shown in Table 4. It can be seen that the scheduling maintenance scheme of the present invention can achieve the lowest cost compared with other schemes, that is, the method provided in Example 1 has feasibility and superiority. And draw as shown Figure 2The scheduling and maintenance Gantt chart shown in the figure shows that the green part is the scheduling agent adaptively selecting the workpiece for processing on the machine, the red part is IM maintenance, and the blue part is CM maintenance. It can be seen that the decision-making provided by the present invention can reduce the maintenance of the machine, so that the machine can be in normal operation for a long time. On this basis, a prediction curve of the remaining service life of machine 1 is drawn as an example. Figure 3 As shown, compare it with Figure 2 From the comparative observation, it can be seen that the maintenance action is a decision made by the maintenance agent adaptively according to the health status of the machine.

[0110] Table 4 Experimental results

[0111] Evaluation Metrics SPT_MT4 EDD_MT4 LWKR_MT4 CR_MT4 Method of the present invention Total cost f 3012 3158 3258 3155 2328 <![CDATA[Delay cost CT T > 1174 1337 1504 1299 960 <![CDATA[Maintenance cost CT M > 1837 1821 1753 1855 1368

[0112] In addition, the present invention also provides a production scheduling and machine maintenance real-time decision-making system based on deep reinforcement learning, which includes a central processing unit (CPU), which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) or loaded from a storage unit into a random access memory (RAM). In RAM, various programs and data required for device operation can also be stored. The CPU, ROM, and RAM are connected to each other via a bus. The input / output (I / O) interface is also connected to the bus.

[0113] Many components in a device are connected to the I / O interface, including: input units, such as a keyboard and mouse; output units, such as various types of displays and speakers; storage units, such as magnetic disks and optical disks; and communication units, such as network cards, modems, and wireless communication transceivers. The communication unit allows the device to exchange information / data with other devices via computer networks such as the Internet and / or various telecommunication networks.

[0114] The processing unit performs the various methods and processes described above, such as methods S1 to S3. For example, in some embodiments, methods S1 to S3 may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program may be loaded and / or installed on the device via a ROM and / or a communication unit. When the computer program is loaded into the RAM and executed by the CPU, one or more steps of methods S1 to S3 described above may be performed. Alternatively, in other embodiments, the CPU may be configured to execute methods S1 to S3 by any other appropriate means (e.g., by means of firmware).

[0115] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0116] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0117] In the context of the present invention, machine-readable medium can be a tangible medium that can contain or store a program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0118] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and such modifications or substitutions are intended to be within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.

Claims

1. A real-time decision-making method for production scheduling and machine maintenance based on deep reinforcement learning, characterized in that: The method performs the following steps on a production workshop where workpieces arrive dynamically, generating real-time decisions on production scheduling and machine maintenance for the production workshop, including: Obtaining the operating characteristics of the production workshop and the actual maintenance status information of the machines, and constructing an integrated optimization model for dynamic scheduling and predictive maintenance based on the operating characteristics and actual maintenance status information; the integrated optimization model includes a cost optimization function and constraints; Acquiring real-time status data of the machine operation, predicting the remaining useful life of the machine using a deep learning method based on the real-time status data, setting a first threshold and a second threshold for the remaining useful life, and classifying the machine maintenance status based on the first threshold and the second threshold; A DDQN-based intelligent agent is constructed, and the integrated optimization model is solved based on the intelligent agent to generate real-time decisions for production scheduling and machine maintenance; the intelligent agent includes a scheduling intelligent agent and a maintenance intelligent agent, and the state of the intelligent agent includes operating characteristics and remaining service life.

2. The method for real-time decision-making for production scheduling and machine maintenance based on deep reinforcement learning according to claim 1, characterized in that: The cost optimization function expression is: min:f=CY T +CY M , Among them, CT T represents the total delay cost of processing the workpiece, and T k represents the delay time of the workpiece, l represents the number of workpieces to be processed, T k =max(0,(C k -D k )) represents workpiece J k The delay time, C k Indicates the workpiece completion time, D k Indicates the delivery time of the workpiece; CT M represents the total maintenance cost of the machine, and CT M =CT CM +CT IM , represents the corrective maintenance cost of the machine, n represents the total number of machines in the production workshop, h i Indicates machine M i Number of corrective maintenance Indicates machine M i Corrective maintenance costs; represents the imperfect maintenance cost of the machine, u i Indicates machine M i The number of imperfect maintenance, w i Indicates the number of machines M per unit time i The cost of imperfect maintenance, Indicates non-perfect maintenance time.

3. The method for real-time decision-making for production scheduling and machine maintenance based on deep reinforcement learning according to claim 2, characterized in that: The calculation method of the non-perfect maintenance time is: Among them, b i represents the linear coefficient, a i Indicates machine M i Basic maintenance time, T x Indicates the running time at which the remaining service life of the machine reaches the first threshold, T t Indicates the current time.

4. The method for real-time decision-making for production scheduling and machine maintenance based on deep reinforcement learning according to claim 1, characterized in that: The constraints include: The completion time of each process of the workpiece needs to be earlier than the start time, the expression is: C kmi >S kmi , where C kmi Indicates workpiece J k The mth process O km On machine M i Completion time, S kmi Indicates process O km On machine M i Start time on The workpiece is processed according to the process sequence, and the expression is: S kmi >C k(m-1)i′ , C k(m-1)i′ Indicates workpiece J k The m-1th process O k(m-1) On machine M i′ Completion time on The maintenance of each machine is only carried out before or after the workpiece is processed. The expression is: (MS i >C kmi )&(ME i k′mi ), MS i Indicates machine M i Maintenance start time, ME i Indicates machine M i Maintenance end time; C kmi Indicates workpiece J k The mth process O km On machine M i Completion time on k′mi Indicates workpiece J k′ The mth process O k′m On machine M i Start time on​ The processing of each step is a continuous process, and the expression is: C kmi =S kmi +p kmi , p kmi Indicates process O km On machine M i Processing time on Each process can only be performed on one machine, and the expression is: X kmi Indicates process O km Assigned to machine M i The status parameter of process O is 1. km Assigned to machine M i , otherwise its value is 0; Each machine only processes one process in the same time period, and the expression is: d k Indicates workpiece J k The number of processes; l represents the number of workpieces to be processed; Each time the machine performs only one task, which includes maintenance and workpiece processing. The maintenance includes imperfect maintenance and corrective maintenance. The expression is: ∑ i∈n (X kmi +Q i +Z i )≤1,X kmi Indicates process O km Assigned to machine M i State parameter; Q i Indicates the state where the machine performs non-perfect maintenance. When Q i =1 indicates that imperfect maintenance is performed, otherwise Q i =0; Z i Indicates the state of the machine performing corrective maintenance. i =1 indicates that imperfect maintenance is performed, otherwise Z i =0; Only when the workpiece reaches the corresponding machine, the process is executed, and the expression is: (C k1i -p k1i -A k )X k1i ≥0, A k Indicates the time when the workpiece arrives at the corresponding machine; p k1i Indicates the processing time of the first step of the workpiece; C k1i Indicates the completion time of the first process of the workpiece; Each machine cannot handle multiple different types of processes at the same time. The expression is: (S kmi ,C kmi )∩(S k′m′i ,C k′m′i )=φ,S k′m′i and C k′m′i Represents workpiece J k′ The m′th process O k′m′ Processing start and end time; S kmi and C kmi Represents workpiece J k The mth process O km The processing start and end time.

5. The method for real-time decision-making for production scheduling and machine maintenance based on deep reinforcement learning according to claim 1, characterized in that: The state space of the scheduling agent is constructed based on the operating characteristics, including the workpiece processing queue operation state and the process operation state; its action space includes: the shortest workpiece process time, the shortest time required to complete the remaining process, the minimum progress ratio of the workpiece processing process and the earliest delivery time; its reward is calculated based on the delay time of the workpiece and the time it is in the process queue state.

6. The method for real-time decision-making for production scheduling and machine maintenance based on deep reinforcement learning according to claim 5, characterized in that: The workpiece processing queue operation status includes: the average difference and minimum difference between the current time when the workpiece is in the queue operation and the delivery time, the total time, average time, minimum time and maximum time required for the remaining processes when the workpiece is in the queue operation, the total time, average time, minimum time and maximum time required for the next process when the workpiece is in the queue operation, the ratio of the maximum time when the workpiece is in the queue operation to the average time when the workpiece is in the process processing, the ratio of the minimum time when the workpiece is in the queue operation to the average time when the workpiece is in the process processing, and the maximum slack time, minimum slack time, average slack time and the sum of the slack time when the workpiece is in the queue operation; The process operation status includes: the number of workpieces being processed, the number of workpieces in queue, the average completion rate of the process processing and the lateness rate of the process processing.

7. The method for real-time decision-making for production scheduling and machine maintenance based on deep reinforcement learning according to claim 5, characterized in that: The reward calculation expression of the scheduling agent is: Among them, R km represents the reward of the scheduling agent; T k Indicates the delay time of the workpiece; AT k Indicates the average delay time of each process; Q km Indicates that the workpiece enters the mth process O km Waiting time; r k Indicates workpiece J k The number of processes.

8. The method for real-time decision-making for production scheduling and machine maintenance based on deep reinforcement learning according to claim 1, characterized in that: The state space of the maintenance agent includes the quantity characteristics, time characteristics and remaining service life of the machine maintenance; its action space is the state of whether the machine is maintained, including no maintenance required, imperfect maintenance and corrective maintenance; its reward is calculated based on the maintenance cost per unit normal operation time between the current maintenance and the last maintenance.

9. The method for real-time decision-making for production scheduling and machine maintenance based on deep reinforcement learning according to claim 8, characterized in that: The quantity characteristics include: the number of corrective maintenance, the number of imperfect maintenance between two consecutive corrective maintenance; The time characteristics include: the interval time between the current time and the previous maintenance, the first difference time between the current remaining service life of the machine and the first threshold, the second difference time between the current remaining service life of the machine and the second threshold, and the difference between the current time and the first difference time.

10. The method for real-time decision-making for production scheduling and machine maintenance based on deep reinforcement learning according to claim 8, characterized in that: The reward calculation method for the maintenance agent is: Among them, r u represents the reward for maintaining the agent; a u represents the action of maintaining the u-th decision point of the agent; DN means no maintenance action is required; CT u represents the maintenance cost at each decision point between two adjacent corrective maintenance tasks; s Indicates the time difference between two adjacent non-DN maintenance times; IT u Represents the time interval between two consecutive maintenance decision points.

Citation Information

Patent Citations

  • Scheduling and maintenance optimization method and system based on embedded reinforcement learning

    CN118735200A

Cited By

  • Automobile door hinge production line dynamic scheduling method and system based on digital twinning

    CN121936869A