An Industrial Internet Preventive Maintenance Method and Device Based on Reinforcement Learning

Through the method based on reinforcement learning, equipment operation parameters and health levels are obtained, and pre-maintenance strategies are optimized using reinforcement learning models, which solves the problem of assembly line equipment maintenance, improves production efficiency and reduces maintenance costs.

CN115291571BActive Publication Date: 2025-07-04BEIJING UNIV OF POSTS & TELECOMM +3
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210738025.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-27
Publication Date
2025-07-04
Estimated Expiration
2042-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively determine the pre-maintenance needs of assembly line equipment in the industrial Internet, resulting in equipment being shut down due to neglected maintenance, increasing maintenance costs and reducing production efficiency.

Method used

Using reinforcement learning-based methods, we use reinforcement learning models to obtain the operating parameters and health levels of the equipment, build vectors, use reinforcement learning models to output pre-maintenance actions, calculate loss costs and reward functions, update model parameters, and optimize pre-maintenance strategies to reduce downtime probability and maintenance costs.

Benefits of technology

It improves the production efficiency of assembly line equipment, reduces the probability of downtime and maintenance costs, and achieves more efficient preventive maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115291571B_ABST
    Figure CN115291571B_ABST
Patent Text Reader

Abstract

The present invention provides an industrial Internet preventive maintenance method and device based on reinforcement learning. The steps of the method include: obtaining the operating parameters of each device on the production line; respectively establishing vectors based on the number of upstream workpieces of each device on the production line and the current health level, and splicing the two vectors to obtain an initial vector; inputting the initial vector into the evaluation network of a preset reinforcement learning model, and the reinforcement learning model outputs the actions to be performed; obtaining the operating parameters of each device on the production line after completing the actions, and establishing an updated vector based on the operating parameters; based on the health conditions of each device on the production line after completing the actions, calculating the loss cost based on the health conditions, and calculating the reward function based on the loss cost; obtaining the target value based on the calculated reward function value and the updated vector, calculating the loss function based on the target value, and updating the parameters of each layer of the reinforcement learning model based on the loss function value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of equipment maintenance, and in particular, to a preventive maintenance method and device for industrial Internet based on reinforcement learning. Background Art

[0002] Currently, the global economic and social development is facing new challenges and opportunities. As a new industrial ecosystem, key infrastructure, and new application model, the industrial Internet is promoting the acceleration of the transformation and upgrading of traditional industries and the development and growth of emerging industries. Equipment fault prediction and health management, as an important part of the industrial Internet, are of great significance for improving the intelligent management level of industrial equipment in national pillar industries such as wind power and national defense.

[0003] In the industrial Internet, for equipment fault prediction and health management, the status monitoring is mainly completed through networked online equipment to realize the transmission and data storage and analysis of equipment status, so as to further predict and determine the damage degree and risk level of the current equipment, and ensure the safe and stable operation of industrial equipment. Through intelligent diagnostic analysis, the networked online monitoring system can not only provide monitoring for the operating status of industrial equipment, but also provide a scientific basis for equipment maintenance, greatly saving the cost in the industrial production process.

[0004] However, for the entire production line, it is difficult to determine which equipment needs to be pre-maintained, which may easily lead to equipment downtime due to lack of maintenance, increased maintenance costs, and reduced production efficiency. Summary of the Invention

[0005] In view of this, embodiments of the present invention provide a preventive maintenance method for industrial Internet based on reinforcement learning to eliminate or improve one or more defects existing in the prior art.

[0006] One aspect of the present invention provides a preventive maintenance method for industrial Internet based on reinforcement learning. The steps of the method include:

[0007] Obtain the operating parameters of each device on the production line, where the operating parameters include the number of upstream workpieces of the device and the current health level of the device;

[0008] Based on the number of upstream workpieces and the current health level of each device on the production line, establish vectors respectively, and splice the two vectors to obtain an initial vector;

[0009] Input the initial vector into the evaluation network of a preset reinforcement learning model, and the reinforcement learning model outputs the action to be taken, where the action is to perform pre-maintenance on a certain device or not to perform pre-maintenance on any device;

[0010] Obtain the operating parameters of each device on the production line after the action is completed, and establish an updated vector based on the operating parameters;

[0011] Based on the health conditions of each device after the pipeline completes an action, calculate the loss cost based on the health conditions, and calculate the reward function based on the loss cost;

[0012] Obtain the target value based on the calculated reward function value and the update vector, calculate the loss function based on the target value, and update the parameters of each layer of the evaluation network of the reinforcement learning model using gradient descent based on the loss function value.

[0013] Adopting the above solution, this solution can output the devices that need preventive maintenance based on the evaluation network of the preset reinforcement learning model. Further, in actual applications, this solution can also, during application, obtain the state of the pipeline after completing an action according to whether preventive maintenance is performed on the pipeline or not. Based on the state of the pipeline after completing the action, the reinforcement learning model can be further updated, and then the maintenance cost can be minimized based on the reward function, thereby reducing the downtime probability and improving production efficiency.

[0014] In some embodiments of the present invention, when inputting the initial vector into the evaluation network of the preset reinforcement learning model, in the step where the reinforcement learning model outputs the action that needs to be performed:

[0015] The evaluation network obtains the initial evaluation value Q1 based on the input initial vector, and outputs the first action corresponding to the maximum value function of the current initial evaluation value Q1.

[0016] In some embodiments of the present invention, the step of obtaining the target value based on the calculated reward function value and the update vector includes:

[0017] Input the update vector into the evaluation network, output the updated evaluation value Q2, output the second action corresponding to the maximum value function of the current initial evaluation value Q2, and obtain the target value based on the second action and the reward function value.

[0018] In some embodiments of the present invention, the step of obtaining the target value based on the second action and the reward function value further includes:

[0019] Input the update vector into the target network of the reinforcement learning model to obtain the evaluation value Q3 output by the target network, and obtain the target value based on the evaluation value Q3 and the reward function value.

[0020] In some embodiments of the present invention, according to the following formula, obtain the target value based on the evaluation value Q3 and the reward function value:

[0021] Q * = R + γ·Q3;

[0022] Q *It represents the target value; R represents the reward function value; γ represents the discount factor; Q3 represents the evaluation value.

[0023] In some embodiments of the present invention, based on the health conditions of each device after the pipeline completes an action, the loss cost is calculated based on the health conditions. The health conditions of the device include whether the device is faulty. The types of the loss cost include pre-maintenance cost, downtime cost, and fault repair cost.

[0024] In some embodiments of the present invention, according to the following formula, the reward function is calculated based on the loss cost:

[0025] R = -(c CBM + k·δ·c CM + c PL );

[0026] R represents the reward function value, c CBM represents the pre-maintenance cost, c CM represents the fault repair cost, c PL represents the downtime cost, k represents the number of faulty devices, and δ represents the penalty factor.

[0027] In some embodiments of the present invention, when the first action is not to perform pre-maintenance on any device, c CBM is 0, and the magnitude of the penalty factor is less than the magnitude of the penalty factor when the first action is to perform pre-maintenance on a certain device.

[0028] In some embodiments of the present invention, the health levels of the devices are preset to multiple. After each action is completed, the health levels of the devices are updated. The update attenuation of the device health levels is a constant relationship, a linear relationship, or a logarithmic relationship;

[0029] If it is a constant relationship, it is shown by the following formula:

[0030] f1(s) = c; s ∈ {s0, s1, s2,..., s N};

[0031] {s0, s1, s2,..., s N} is the set of health levels, s represents the health level, c is the health level attenuation probability value corresponding to s, and f1(s) is the health level attenuation probability corresponding to s under the constant relationship;

[0032] If it is a linear relationship, it is shown by the following formula:

[0033] f2(s) = ks + b; s ∈ {s0, s1, s2,..., s N ,};

[0034] Both k and b are preset parameters, s represents the health level, and f2(s) is the health level decay probability corresponding to s under the linear relationship;

[0035] If it is a logarithmic relationship, it is shown by the following formula:

[0036] f3(s) = logas + d; s ∈ {s0, s1, s2,..., s N ,};

[0037] Both d and a are preset parameters, and f3(s) is the health level decay probability corresponding to s under the logarithmic relationship.

[0038] In some embodiments of the present invention, there is a failure probability corresponding to each health level of the device, and the preset failure probability follows the Weibull distribution.

[0039] In some embodiments of the present invention, according to the following formula, the loss function is calculated based on the target value:

[0040] Loss = (Q * -Q1) 2 ;

[0041] Loss represents the value of the loss function, Q * represents the target value, and Q1 represents the initial evaluation value.

[0042] Another aspect of the present invention also provides an industrial Internet preventive maintenance device based on reinforcement learning. The device includes a computer device, the computer device includes a processor and a memory, the memory stores computer instructions, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the device implements the steps implemented by the method described above.

[0043] The additional advantages, objectives, and features of the present invention will be partially described below, and will become partially obvious to those of ordinary skill in the art after studying the following text, or can be learned from the practice of the present invention. The objectives and other advantages of the present invention can be pointed out and obtained specifically in the description and the accompanying drawings.

[0044] Those skilled in the art will understand that the objectives and advantages that can be achieved by the present invention are not limited to the above specifically described, and the above and other objectives that the present invention can achieve will be more clearly understood according to the following detailed description. Description of the Drawings

[0045] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not constitute a limitation to the present invention.

[0046] Figure 1Schematic diagram of an implementation manner of the preventive maintenance method for industrial Internet based on reinforcement learning of the present invention;

[0047] Figure 2 Schematic diagram of the structure of the production line of the present invention;

[0048] Figure 3 Schematic diagram of the state transition process of the production line equipment;

[0049] Figure 4 Graph of comparative analysis of maintenance costs in the experimental example;

[0050] Figure 5 Graph of comparison of initial production in the experimental example;

[0051] Figure 6 Graph of comparison of final production in the experimental example. Detailed implementation manner

[0052] To make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below in combination with the implementation manners and the accompanying drawings. Herein, the illustrative implementation manners of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention.

[0053] Herein, it also needs to be noted that in order to avoid obscuring the present invention due to unnecessary details, only the structures and / or processing steps closely related to the solution according to the present invention are shown in the drawings, while other details less related to the present invention are omitted.

[0054] It should be emphasized that the term "including / comprising" when used herein refers to the presence of features, elements, steps or components, but does not exclude the presence or addition of one or more other features, elements, steps or components.

[0055] Herein, it also needs to be noted that if not otherwise specified, the term "connection" in this text can not only refer to direct connection, but also represent indirect connection with an intermediate.

[0056] In the following, embodiments of the present invention will be described with reference to the accompanying drawings. In the drawings, the same reference numerals represent the same or similar components, or the same or similar steps.

[0057] Introduction to the prior art:

[0058] Prior art 1 proposed an optimization method for condition-based maintenance strategy of accelerated degradation equipment based on Monte Carlo simulation. This method determines the characteristics of equipment accelerated degradation based on the Gamma stochastic process and obtains the average cost rate model under the long-term operation conditions of the equipment by determining the influence of the degradation state of the equipment on product quality.

[0059] The prior art 2 proposes a preventive maintenance method based on a continuous-time MDP model, which comprehensively considers various factors affecting the production line efficiency and uses the continuous-time MDP model to simulate the equipment degradation process.

[0060] The prior art 3 proposes a predictive maintenance method and system for manufacturing equipment. The method of reinforcement learning is used to determine the pre-maintenance strategy corresponding to each equipment state, and a correspondence table between the equipment state and the pre-maintenance strategy is obtained; finally, according to the current equipment state, the pre-maintenance strategy corresponding to the current equipment state is determined.

[0061] The prior art 1 is suitable for scenarios where the equipment state is closely related to the product qualification rate and has certain constraints. The present invention aims at maximizing output and minimizing maintenance costs and has a wider application scenario. The prior art 2 only proposes a preventive maintenance strategy in terms of framework and does not analyze the specific state transition process of different equipment. The present invention considers the maintenance schemes for three specific state decay rate devices and has better universality. The prior art 3 uses a single strong network, and the maximization operation in the network update process leads to the estimated Q value being larger than the true value. The present invention selects to use the DDQN algorithm to implement the agent, which can effectively prevent this problem.

[0062] To solve the above problems, as Figure 1 shown, the present invention proposes a preventive maintenance method for industrial Internet based on reinforcement learning. The steps of the method include:

[0063] Step S100, obtaining the operation parameters of each device on the production line, where the operation parameters include the number of upstream workpieces of the device and the current health level of the device.

[0064] In some embodiments of the present invention, a buffer for carrying workpieces to be processed is provided upstream of each device on the production line, and the number of upstream workpieces of the device is the number of workpieces to be processed in the buffer corresponding to the device.

[0065] The production line system of the present invention consists of N different devices M = {m1, m2,..., m N} and intermediate buffers B = {b2, b3,..., b N}, as Figure 2 shown. In the production line system, the components processed by the device will be put into the downstream buffer and wait for the downstream device to process it. Therefore, different devices are inseparable.

[0066] The state of each device can be described by a triple, namely {t, s, b'}. Among them, t represents the running time of the device; s represents the current device health level, and the initial state of each new device during model training is 0, indicating that the device is in the best state; b' represents the size of the upstream buffer of the current device. The production system is modeled as a discrete event and time simulation. When the device is running, it obtains the production components of the upstream device from the upstream buffer and starts to process it. During the processing, these devices may degrade. Since the probability d of the device decaying from the current state to the next state is uncertain, this paper formulates the decay process as a discrete semi-Markov process. The difference between the semi-Markov process and the Markov process lies in that its transition time and probability depend on the time the system is in the current state.

[0067] It should be noted that during the use of the device, in addition to the device decay state, there is also a failure state. In the device decay state, the device can still produce, but the failure state means that the device cannot work. The relationship between the device decay state and the failure state is as Figure 3 shown. Figure 3 shows the state transition process of the pipeline device, where the device has N i +1 states. When the device is in state N i , it means that it is in the failure state. The transition probability of the decay state is determined by the nature of the device itself. For different devices, its decay process, that is, the semi-Markov process, is different. Before the device reaches the failure state, the device can take condition-based maintenance (CBM). If the device is already in the failure state (N i ), then only corrective maintenance (CM) can be performed to make the device work again.

[0068] In some embodiments of the present invention, each device's health level is preset with multiple levels, and the lower the health level, the higher the failure probability.

[0069] Step S200, respectively establish vectors based on the number of upstream workpieces and the current health level of each device on the pipeline, and splice the two vectors to obtain an initial vector;

[0070] In some embodiments of the present invention, if there are three devices on the pipeline, the number of upstream workpieces of the devices are 5, 10, and 7 respectively; the health levels of the devices are 1, 2, and 3 respectively. Then the vectors (5, 10, 7) and (1, 2, 3) are established respectively, and splicing the two vectors gives the initial vector (5, 10, 7, 1, 2, 3).

[0071] Step S300: Input the initial vector into the evaluation network of a preset reinforcement learning model. The reinforcement learning model outputs the action to be taken, which is to perform pre-maintenance on a certain device or not to perform pre-maintenance on any device.

[0072] In some embodiments of the present invention, the reinforcement learning model is a DDQN model. The reinforcement learning model includes an evaluation network and a target network. During the training process of the reinforcement learning model, it can be preset that every n rounds of training, the parameters of each layer of the target network are replaced with the parameters of each layer of the evaluation network.

[0073] In some embodiments of the present invention, the action output by the reinforcement learning model in this solution can be an instruction to perform pre-maintenance on a device with a specified number, or not to perform pre-maintenance on any device in the pipeline. The pre-maintenance can be in the form of software maintenance or hardware maintenance.

[0074] Step S400: Obtain the operating parameters of each device in the pipeline after the action is completed, and establish an update vector based on the operating parameters.

[0075] In some embodiments of the present invention, the method of establishing the update vector based on the operating parameters in this step is the same as the method of obtaining the initial vector in step S200.

[0076] In some embodiments of the present invention, the operating parameters of the devices in the pipeline are updated after the action is completed, that is, it may include reducing or increasing their own health levels.

[0077] The increase in the health level can be based on the pre-maintenance action.

[0078] Step S500: Calculate the loss cost based on the health conditions of each device in the pipeline after the action is completed, and calculate the reward function based on the loss cost.

[0079] In some embodiments of the present invention, the health conditions of each device can be that none of the devices have failed or one or more of the devices have failed. If a device fails, the losses include the repair cost of the downtime failure and the downtime cost. If the device was pre-maintained in the previous round, the losses also include the pre-maintenance cost.

[0080] If the device was pre-maintained in the previous round and did not fail, the losses include the pre-maintenance cost.

[0081] Step S600: Obtain the target value based on the calculated reward function value and the update vector, calculate the loss function based on the target value, and update the parameters of each layer of the evaluation network of the reinforcement learning model using gradient descent based on the loss function value.

[0082] In some embodiments of the present invention, the present solution establishes a reward function based on the loss cost, enabling the model to perform actions tending towards the reward, i.e., actions with lower losses, and minimizing the loss on the premise of improving production efficiency.

[0083] Adopting the above solution, the present solution can output the devices that need to be pre-maintained based on the evaluation network of the preset reinforcement learning model. Further, in practical applications, the present solution can also, during application, obtain the state of the production line after the completion of the action based on the action of pre-maintaining or not pre-maintaining the production line according to the production line. Based on the state of the production line after the completion of the action, the reinforcement learning model can be further updated, and then the maintenance cost can be minimized based on the reward function, thereby reducing the downtime probability and improving production efficiency.

[0084] In some embodiments of the present invention, in the step of inputting the initial vector into the evaluation network of the preset reinforcement learning model and the reinforcement learning model outputting the actions that need to be performed:

[0085] The evaluation network obtains an initial evaluation value Q1 based on the input initial vector, and outputs the corresponding first action for which the current initial evaluation value Q1 obtains the maximum value function.

[0086] In some embodiments of the present invention, the evaluation network obtains an initial evaluation value Q1 based on the input initial vector, which can be expressed by the following formula:

[0087] a1 = argmaxQ1;

[0088] a1 represents the first action, where Q1 can be expressed as Q estimate (s t ,a t ; θ), Q estimate represents the evaluation network, a t represents the action output by the evaluation network in the previous round of this round, θ represents the parameters of each layer of the current evaluation network, and s t represents the operating parameters of each device of the current production line.

[0089] In some embodiments of the present invention, the step of obtaining the target value based on the calculated reward function value and the update vector includes:

[0090] Input the update vector into the evaluation network, output the updated evaluation value Q2, output the corresponding second action for which the current initial evaluation value Q2 obtains the maximum value function, and obtain the target value based on the second action and the reward function value.

[0091] In some embodiments of the present invention, the output of the corresponding second action for which the current initial evaluation value Q2 obtains the maximum value function can be expressed as the following formula:

[0092] a2 = argmax Q2;

[0093] a2 represents the second action, where Q2 can be expressed as Q estimate (s t +1, a t+1 ; θ), Q estimate represents the evaluation network, a t+1 represents the action output by the evaluation network in this round, θ represents the parameters of each layer of the current evaluation network, s t+1 represents the operating parameters of each device in the pipeline after the first action.

[0094] In some embodiments of the present invention, the step of obtaining the target value based on the second action and the reward function value further includes:

[0095] Input the update vector into the target network of the reinforcement learning model to obtain the evaluation value Q3 output by the target network, and obtain the target value based on the evaluation value Q3 and the reward function value.

[0096] In some embodiments of the present invention, input the update vector into the target network of the reinforcement learning model to obtain the evaluation value Q3 output by the target network, and Q3 can be expressed by the following formula:

[0097] Q target (s t+1 , a2; θ');

[0098] Q target represents the target network, s t+1 represents the operating parameters of each device in the pipeline after the first action, a2 represents the second action, and θ' represents the parameters of each layer of the current target network.

[0099] In some embodiments of the present invention, according to the following formula, obtain the target value based on the evaluation value Q3 and the reward function value:

[0100] Q * = R + γ·Q3;

[0101] Q * represents the target value; R represents the reward function value; γ represents the discount factor; Q3 represents the evaluation value.

[0102] In some embodiments of the present invention, based on the health conditions of each device in the pipeline after completing the action, calculate the loss cost based on the health conditions. The health conditions of the device include whether the device is faulty, and the types of the loss cost include pre-maintenance cost, downtime cost, and fault repair cost.

[0103] In some embodiments of the present invention, the pre-maintenance cost, downtime cost, and fault repair cost can all be preset time cost parameters or resource consumption parameters, and the consumed resources can be money, etc.

[0104] In some embodiments of the present invention, the calculation formula of the downtime cost at time step t is as follows:

[0105]

[0106] c PL (t) represents the downtime cost at time step t, and T total represents the total downtime duration, and T M' represents the duration required for the device with the slowest workpiece processing speed in the pipeline to process one workpiece.

[0107] Adopting the above scheme, the loss of the downtime cost is determined based on the total downtime duration and the duration required for the device with the slowest workpiece processing speed in the pipeline to process one workpiece, improving the calculation accuracy of the downtime cost.

[0108] In some embodiments of the present invention, the goal of the maintenance plan is to ensure a high equipment availability while using as few maintenance resources as possible. Considering that ensuring a high level of availability of the production system incurs a high cost of regularly performing maintenance operations, it often conflicts with the short-term operation goal of achieving low costs during production. Considering that the production line involving maintenance activities is highly complex and random, with a large state space and difficult to handle, traditional planning methods are difficult to obtain an optimal maintenance strategy. However, reinforcement learning has natural advantages in dealing with such problems. Reinforcement learning is particularly suitable for problems including long-term and short-term return trade-offs. In a production scenario, CBM is increasing short-term costs. However, it will save long-term costs by avoiding potential CM.

[0109] In some embodiments of the present invention, according to the following formula, the reward function is calculated based on the loss cost:

[0110] R = -(c CBM + k·δ·c CM + c PL );

[0111] R represents the value of the reward function, c CBM represents the pre-maintenance cost, c CM represents the breakdown maintenance cost, c PL represents the downtime cost, k represents the number of faulty devices, and δ represents the penalty factor.

[0112] In some embodiments of the present invention, when the first action is not to perform pre-maintenance on any device, c CBM is 0, and the magnitude of the penalty factor is less than the magnitude of the penalty factor when the first action is to perform pre-maintenance on a certain device.

[0113] In some embodiments of the present invention, when the first action is not to perform preventive maintenance on any device, the reward function is calculated based on the loss cost according to the following formula:

[0114] R = -(k·α·c CM + c PL );

[0115] When the first action is to perform preventive maintenance on any device, the reward function is calculated based on the loss cost according to the following formula:

[0116] R = -(c CBM + k·β·c CM + c PL );

[0117] R represents the reward function value, c CBM represents the preventive maintenance cost, c CM represents the fault repair cost, c PL represents the downtime cost, k represents the number of faulty devices, and both α and β represent penalty factors, where α > β.

[0118] Adopting the above scheme, if the device fails without preventive maintenance, it indicates that the evaluation network makes inaccurate actions. Therefore, the penalty factor α > β is set to improve the efficiency of model training and the accuracy of the output actions.

[0119] From the composition of the reward function, it can be seen that the setting of the reward function is to minimize the maintenance cost. Among them, the reward function needs to take the negative of these costs to maximize the reward. In this way, a large-scale pipeline system maintenance model based on reinforcement learning is defined and can be solved by DDQN and a data-driven modeling simulator.

[0120] As a general way to describe the device decay process, the Markov chain is applicable to describe devices with discrete and continuously decaying states. The continuous device states can be discretized into a finite number of independent states, and the sojourn time of the states follows a certain distribution. Since the production maintenance cost is always related to some time frames, more problems in actual production tend to be described as Markov processes or semi-Markov processes. Compared with the Markov process, in the semi-Markov process, the device state transition time and probability depend on the system reaching the current state. Therefore, it can better describe the device state and the model has higher accuracy. So the present invention tends to describe the device decay process as a semi-Markov decision process.

[0121] In a semi-Markov process, the sojourn time distribution of a state can adopt any positive random variable. Since there are numerous system functions in the industrial Internet and the corresponding device characteristics are also different. Therefore, the present invention applies a maintenance model to devices with different state decay processes, so as to analyze the effects of the model in different processes and whether the model has universality. The present invention assumes that the device state S = {s0, s1, s2,..., s N}, and the functional relationship distribution between the device state and the failure rate includes a constant relationship, a linear relationship, or a logarithmic relationship.

[0122] In some embodiments of the present invention, the health levels of the devices are preset to be multiple, and the health levels of the devices are updated after each action is completed. The update decay of the device health levels is a constant relationship, a linear relationship, or a logarithmic relationship;

[0123] If it is a constant relationship, it is shown by the following formula:

[0124] f1(s) = cs ∈ {s0, s1, s2,..., s N};

[0125] {s0, s1, s2,..., s N} is the set of health levels, s represents the health level, c is the health level decay probability value corresponding to s, and f1(s) is the health level decay probability corresponding to s under the constant relationship;

[0126] If it is a linear relationship, it is shown by the following formula:

[0127] f2(s) = ks + b s ∈ {s0, s1, s2,..., s N ,};

[0128] Both k and b are preset parameters, s represents the health level, and f2(s) is the health level decay probability corresponding to s under the linear relationship;

[0129] If it is a logarithmic relationship, it is shown by the following formula:

[0130] f3(s) = log a s + d s ∈ {s0, s1, s2,..., s N ,};

[0131] Both d and a are preset parameters, and f3(s) is the health level decay probability corresponding to s under the logarithmic relationship.

[0132] By adopting the above scheme, there is a certain probability of health level decay after each action is taken. By adopting the above method, the predictability of the decay is improved.

[0133] In some embodiments of the present invention, there is a corresponding failure probability for each health level of the device, and the preset failure probability follows a Weibull distribution.

[0134] In some embodiments of the present invention, the loss function is calculated based on the target value according to the following formula:

[0135] Loss=(Q * -Q1) 2 ;

[0136] Loss represents the loss function value, Q * represents the target value, and Q1 represents the initial evaluation value.

[0137] In an industrial Internet production system, the state of equipment deteriorates continuously due to fatigue, wear, aging, and various uncertainties in the dynamic environment, resulting in a gradual decline in work efficiency and performance. If not repaired, it will ultimately be unable to meet the work requirements and even cause serious failures. The maintenance costs and downtime losses caused by equipment failures are quite huge and can account for 30%-40% of the total production cost in the absence of a reasonable maintenance strategy. Although many scholars at home and abroad have conducted extensive research on equipment maintenance problems in production systems from multiple perspectives, the association and combination of equipment maintenance and the latest artificial intelligence in recent years are not in-depth.

[0138] The present invention models multiple devices with different equipment decay rates and their upstream buffers in a pipeline system, proposes a preventive maintenance strategy based on DDQN, obtains the optimal maintenance methods for different types of pipeline equipment under resource constraints, and achieves the purpose of increasing production and reducing maintenance costs, which has a certain guiding role in solving equipment maintenance problems in the actual production process.

[0139] The embodiment of the present invention further provides an industrial Internet preventive maintenance device based on reinforcement learning. The device includes a computer device, the computer device includes a processor and a memory, the memory stores computer instructions, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the device realizes the steps implemented by the method described above.

[0140] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps implemented by the foregoing industrial Internet preventive maintenance method based on reinforcement learning are realized. The computer-readable storage medium may be a tangible storage medium, such as a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, floppy disk, hard disk, removable storage disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0141] For a pipeline system with N devices and N buffers, the overall goal of the maintenance strategy is to optimize the costs associated with equipment maintenance. Generally speaking, maintaining equipment will stop its operation. During maintenance, corrective maintenance lasts longer than condition-based maintenance, which means that corrective maintenance will cause more production losses. From an economic perspective, the maintenance cost consists of two parts: one part is the repair cost for replacing old or faulty equipment; the other part of the cost is related to the production losses caused by equipment downtime. There is a trade-off problem in this decision-making process. If the CBM is insufficient, the total CM will increase, and long-term corrective maintenance will seriously affect the system performance. If the CBM is too frequent, although random failures can be significantly reduced, the increase in CBM costs may far exceed the costs saved by avoiding CM. Our goal is to find an optimal strategy to minimize the cost of equipment maintenance, and the objective function is described as follows:

[0142]

[0143] where π represents the maintenance strategy adopted by the current pipeline system, and are the number of preventive maintenance and corrective maintenance performed on device i, respectively. PL refers to the production loss of the system due to all maintenance activities.

[0144] The beneficial effects of the present invention include:

[0145] 1. The present invention conducts relevant research on preventive maintenance strategies for decaying equipment in a pipeline system, and proposes a pipeline equipment maintenance model based on reinforcement learning. A unified agent completes the allocation of maintenance resources and the selection of maintenance methods for pipeline equipment, improving production while reducing maintenance costs.

[0146] 2. Considering the diversity of system types in the industrial Internet, the present invention describes the decay process of equipment by establishing different semi-Markov decision process models, and on this basis, proposes a reinforcement learning algorithm that can be applied to different decay processes.

[0147] 3. Considering that in the traditional deep Q - network (DQN) model of reinforcement learning, the max operation is used to quickly move the Q - value closer to the possible optimization goal. However, this approach can lead to over - estimation, resulting in a large deviation in the model. In this paper, Double DQN (DDQN) is selected to decouple the two steps of selecting the action of the target Q - value and calculating the target Q - value, so as to eliminate the problem of over - estimation.

[0148] 4. The present invention proposes a pipeline system maintenance model based on reinforcement learning and realizes the system maintenance scheduling through DDQN. The model finally evaluates and makes a horizontal comparison of different equipment decay processes. It can be found that adopting a preventive maintenance scheduling scheme based on reinforcement learning within a fixed period of time can effectively save costs during the maintenance process and improve the overall output. Thus, it can be seen that the invention is universal for conventional equipment in different decay states and can play a certain guiding role in solving the equipment maintenance problems in the actual production process.

[0149] Experimental example:

[0150] In order to demonstrate the applicability of reinforcement learning to streamline systems with different decay processes, the present invention sets up three different semi - Markov decision processes. In Configuration I, the equipment is characterized by a relatively fixed state decay rate in its decay process, and the decay rate c = 0.25; in Configuration II, the decay rate of the equipment at the current state is positively correlated with the current state and has a fixed growth amplitude, where k = 0.05 and b = 0.2; in Configuration III, the decay rate of the equipment at the current state has a logarithmic relationship with the current state, and the growth amplitude of its decay rate gradually slows down, where a = 15 and d = 0.05. Except for the functional relationship between the state decay rate and the current state, the other parameters of the system are the same. The component processing speed v and the upstream buffer capacity b are shown in Table 1. In addition, the system sets the pre - maintenance time t CBM = 5, t CM = 20, the number of equipment health levels is 10, the penalty factor β = 10, and the maintenance costs are c CBM = 0.5, c CM = 1.5, c PL = 0.5;

[0151] Table 1 Production equipment configuration table

[0152] M1 M2 M3 M4 M5 Processing speed (v) 1 2 1 1 2 Buffer size (b) 2 2 2 2 2

[0153] Result analysis

[0154] During the training process, the number of execution rounds for each system is set to 2000 rounds. In each round, the simulation time is set to 400. The maintenance costs of different systems change with the training process as Figure 4 shown. From Figure 4It can be found that Configuration III has stronger volatility during the training process compared to Configuration I and Configuration II. The maintenance costs of Configuration I and Configuration II can reach a steady state earlier than that of Configuration III. Generally speaking, after being trained by the DDQN network, the three systems with different decay rates can all make relatively ideal decisions, significantly reducing the maintenance costs.

[0155] Another indicator to measure the quality of the system is the output. To more intuitively analyze the number of final production components under different systems, Figure 5 、 Figure 6 The number of production components of different systems in the first 100 rounds and the last 100 rounds are respectively intercepted. It can be found from the figure that at the same time, the number of production components after the training of System I and System III is similar, and both are higher than that of System II. By comparing Figure 5 with Figure 6 , it can be found that the total production of the three systems has been greatly improved compared to the random maintenance state before training. It can be seen that the optimal maintenance strategy based on DDQN can effectively increase the output on the basis of reducing the maintenance cost.

[0156] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in combination with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to implement it in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, etc. When implemented in software, the elements of the present invention are programs or code segments used to perform the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave on a transmission medium or a communication link.

[0157] It should be clear that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, the detailed description of known methods is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present invention.

[0158] In the present invention, the features described and / or illustrated for one embodiment can be used in the same or a similar manner in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.

[0159] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various modifications and variations can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An industrial Internet preventive maintenance method based on reinforcement learning, characterized in that, The steps of the method include: Obtain the operating parameters of each device on the production line, where the operating parameters include the number of upstream workpieces of the device and the current health level of the device; Respectively establish vectors based on the number of upstream workpieces and the current health level of each device on the production line, and splice the two vectors to obtain an initial vector; Input the initial vector into the evaluation network of a preset reinforcement learning model, and the reinforcement learning model outputs the action to be taken, where the action is to perform preventive maintenance on a certain device or not to perform preventive maintenance on any device; Obtain the operating parameters of each device on the production line after the action is completed, and establish an updated vector based on the operating parameters; Based on the health conditions of each device on the production line after the action is completed, calculate the loss cost based on the health conditions, and calculate the reward function based on the loss cost; Calculate the reward function based on the loss cost according to the following formula: R = -(c CBM + k·δ·c CM + c PL ); R represents the reward function value, c CBM represents the pre-maintenance cost, c CM represents the breakdown maintenance cost, c PL represents the downtime cost, k represents the number of faulty devices, and δ represents the penalty factor; Obtain the target value based on the calculated reward function value and the updated vector, input the updated vector into the evaluation network, output the updated evaluation value Q2, output the corresponding second action of the current initial evaluation value Q2 to obtain the maximum value function, input the updated vector into the target network of the reinforcement learning model, obtain the evaluation value Q3 output by the target network, obtain the target value based on the evaluation value Q3 and the reward function value, according to the following formula: Q * = R + γ·Q3; Q * represents the target value; R represents the reward function value; γ represents the discount factor; Q3 represents the evaluation value; Calculate the loss function based on the target value, and update the parameters of each layer of the evaluation network of the reinforcement learning model by using the backpropagation method based on the loss function value.

2. The preventive maintenance method for industrial Internet based on reinforcement learning according to claim 1, wherein In the step of inputting the initial vector into the evaluation network of a preset reinforcement learning model, and the reinforcement learning model outputs the action to be taken: The evaluation network obtains the initial evaluation value Q1 based on the input initial vector, and outputs the corresponding first action of the current initial evaluation value Q1 to obtain the maximum value function.

3. The preventive maintenance method for industrial Internet based on reinforcement learning according to claim 2, characterized in that, When the first action is not to perform pre-maintenance on any device, c CBM is 0, and the magnitude of the penalty factor is less than the magnitude of the penalty factor when the first action is to perform pre-maintenance on a certain device.

4. The preventive maintenance method for industrial Internet based on reinforcement learning according to any one of claims 1-3, characterized in that, The health level of the device is preset to be multiple, and the health level of the device is updated after each action is completed. The update attenuation of the device health level is a constant relationship, a linear relationship or a logarithmic relationship; If it is a constant relationship, it is shown in the following formula: f1(s) = c; s ∈ {s0, s1, s2,..., s N}; {s0, s1, s2,..., s N} is the set of health levels, s represents the health level, c is the health level decay probability value corresponding to s, and f1(s) is the health level decay probability corresponding to s under constant relationship; If it is a linear relationship, it is shown in the following formula: f2(s) = ks + b; s ∈ {s0, s1, s2,..., s N ,}; Both k and b are preset parameters, s represents the health level, and f2(s) is the health level attenuation probability corresponding to s under the linear relationship; If it is a logarithmic relationship, it is shown in the following formula: f3(s) = log a s + d; s ∈ {s0, s1, s2,..., s N ,} Both d and a are preset parameters, and f3(s) is the health level attenuation probability corresponding to s under the logarithmic relationship.

5. The preventive maintenance method for industrial Internet based on reinforcement learning according to claim 2, wherein Calculate the loss function based on the target value according to the following formula: Loss=(Q * -Q1) 2 ; Loss represents the loss function value, Q * represents the target value, and Q1 represents the initial evaluation value.

6. An industrial Internet preventive maintenance device based on reinforcement learning, characterized in that, The device includes a computer device, the computer device includes a processor and a memory, the memory stores computer instructions, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the device realizes the steps implemented by the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Robot control method based on offline model pre-training learning DDPG algorithm

    CN112668235A

  • Production and maintenance joint optimization method for serial production system based on reinforcement learning

    CN113112051A