Reinforcement Learning Method and Computer-Readable Medium for Maintenance Decision Making

The offline RL system addresses inefficiencies in predictive maintenance by using past data to directly generate maintenance actions, enhancing automation and scheduling efficiency without relying on simulators, thus optimizing maintenance operations.

JP7708834B2Active Publication Date: 2025-07-15HITACHI LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2023191931
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-01-27
Filing Date
2023-11-10
Publication Date
2025-07-15
Estimated Expiration
2043-11-10

AI Technical Summary

Technical Problem

Existing predictive maintenance methods rely heavily on human interpretation of machine learning outputs, leading to inefficiencies in determining optimal maintenance times and costs, and lack effective automation for complex maintenance scheduling.

Method used

An offline reinforcement learning (RL) system that uses past observations and actions to predict maintenance actions directly, incorporating a decision maker model that generates maintenance decisions without requiring a simulator, and provides explainable AI for user understanding.

Benefits of technology

Enables automated, efficient predictive maintenance decisions with reduced downtime and costs by directly generating optimal actions, reducing the need for high-fidelity simulators and improving maintenance scheduling through data-driven approaches.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007708834000008
    Figure 0007708834000008
  • Figure 0007708834000009
    Figure 0007708834000009
  • Figure 0007708834000010
    Figure 0007708834000010
Patent Text Reader

Abstract

To provide a predictive maintenance method for deciding appropriate timing of maintenance activities that avoid premature or unnecessary maintenance and, at the same time, reduce the risk of equipment downtime associated with unexpected failures.SOLUTION: A method includes: receiving expected future return value as input to a decision maker model, the decision maker model being a machine learning model that predicts maintenance action associated with the equipment; feeding recent observations and recent actions from environment as inputs to the decision maker model; generating a next action as model outputs of the decision maker model, the next action being the predicted maintenance action; and executing the next action in the environment.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to methods, computer-readable media, and systems for using offline reinforcement learning (RL) for predictive maintenance of equipment.

Background Art

[0002] Industrial machinery is subject to aging deterioration. Reduction of maintenance costs is one of the major concerns for industrial organizations. Maintenance costs include not only equipment repair costs and labor costs of workers, but also other factors such as economic losses due to equipment downtime and physical occupancy space during repair and replacement. If the inspection frequency of equipment is reduced to cut repair costs, the probability of sudden failures increases, and once a failure occurs, the replacement cost becomes high. If the inspection and repair frequency of equipment is high, the risk of sudden failures is low, but the maintenance cost increases.

[0003] In related art, past observations of equipment inspected at time intervals are used for failure prediction in subsequent time intervals. However, it is a human operator / user who receives the output of the model and determines the repair time of the equipment.

[0004] In related art, machine learning models (e.g., LSTM and functional neural networks) are used for time series analysis to estimate the remaining useful life (RUL) based on observation and failure history records. Also in this case, it is a human operator / user who receives the output of the model and determines the repair time of the equipment.

[0005] It is necessary to perform smart preventive / predictive maintenance to determine the appropriate timing of maintenance activities that avoid premature or unnecessary maintenance and at the same time reduce the risk of equipment downtime associated with unexpected failures. An effective maintenance plan must also consider various constraints such as the maximum number of equipment that can be repaired simultaneously.

Summary of the Invention

Problems to be Solved by the Invention

[0006] Aspects of the present disclosure include innovative methods for predictive maintenance of equipment. The method includes receiving a future expected return value as an input to a decision maker model, the decision maker model being a machine learning model that predicts maintenance actions related to the equipment. The method further includes supplying recent observations and recent actions from the environment as inputs to the decision maker model and generating a next action as the model output of the decision maker model. This next action is the predicted maintenance action. The method may include performing the next action in the environment.

[0007] Aspects of the present disclosure include a non-transitory computer-readable medium storing instructions for predictive maintenance of equipment. The instructions receive a future return value expected as an input to a decision maker model, the decision maker model being a machine learning model that predicts maintenance actions related to the equipment, and the instructions include supplying recent observations and recent actions from the environment as inputs to the decision maker model and generating a next action as the model output of the decision maker model. This next action is the predicted maintenance action. Further, the instructions may include performing the next action in the environment.

[0008] Aspects of the present disclosure include an innovative server system for predictive maintenance of equipment. This system receives future expected return values as inputs to a decision-making model. This decision-making model is a machine learning model that predicts maintenance actions related to the equipment. The server system supplies recent observations and recent actions from the environment as inputs to the decision-making model. It generates the next action as the model output of this decision-making model. This next action is the predicted maintenance action. The server system may include executing the next action in the environment.

[0009] Aspects of the present disclosure include an innovative system for predictive maintenance of equipment. This system includes means for receiving future expected return values as inputs to a decision-making model. This decision-making model is a machine learning model that predicts maintenance actions related to the equipment. This system includes means for supplying recent observations and recent actions from the environment as inputs to the decision-making model, and means for generating the next action as the model output of the decision-making model. The next action is the predicted maintenance action. It can also include means for executing the next action in the environment.

[0010] Next, a general architecture for implementing various features of the present disclosure will be described with reference to the drawings. The drawings and related descriptions are provided to illustrate exemplary implementations of the present disclosure and do not limit the scope of the present disclosure. Throughout the drawings, reference numerals are reused to indicate the correspondence between the elements being referenced.

Brief Description of the Drawings

[0011]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

[0012] The following detailed description provides details of the figures and embodiments of the present application. References between figures and descriptions of redundant elements are omitted for clarity. The terms used throughout this specification are provided as examples and are not intended to be limiting. For example, the use of the term "automatic" can include fully automatic or semi-automatic embodiments with user or administrator control over specific aspects of the embodiment, depending on the desired embodiment of those skilled in the art practicing the embodiments of the present invention. The selection can be performed by the user through a user interface or other input means, or can be performed through a desired algorithm. The exemplary embodiments described herein can utilize either singular or combinations, and the functions of the exemplary embodiments can be implemented through any means according to the desired embodiment.

[0013] Sequential decision-making under uncertainty is more difficult than one-shot optimization because the current state of the system changes continuously over time and the future state depends non-linearly on the current and past decisions. Deep neural network-based reinforcement learning (RL) is a promising approach to tackle such problems. An RL agent acts in a Markov environment with state st. The next state s t+1 is sampled from the transition probability P(s t , a t ). Here, a t is the action of the RL agent in state s t . The RL agent receives a reward r (st, at) ∈ R at each time step. Here, r is the Greek letter gamma. The goal is to learn a policy "pi" (a map from the state space to the action space) that maximizes the expected total reward E[Sigma tr t r t over a number of time steps called an episode. Such cumulative reward is often called return.

[0014] Figure 6 shows the differences between offline deep RL learning and training in a real system, and training in a simulator. RL attempts to optimize an entire sequence of actions to solve a temporarily extended control problem. There are few successful examples of applying RL to real-world problems due to the required training time and associated costs. Also, in training using a real system, safety in real-time operation is a concern. On the other hand, training on a simulator in a conventional online RL setting is costly and sometimes infeasible. Also, modeling the failure mode itself is an issue in training on a simulator. In contrast to conventional online RL algorithms that require a high-fidelity simulator, offline RL does not require a simulator and instead learns useful skills from a dataset of past actions and observations.

[0015] Accurate simulators of machines are rarely available, but sensor data and repair history records can be easily obtained without additional cost. If the dataset contains little or no repair data, repair data can be artificially inserted during the episode and the device can be reset to its initial state. Offline RL does not require live interaction with the environment and develops a useful set of skills from the dataset. Maintenance optimization can be formulated as a supervised machine learning problem, enabling the RL agent to acquire effective maintenance skills even when the maintenance history record consists of logs of bad operations. Furthermore, the RL agent is equipped with an explainable AI module, which makes it easier for users to interpret the model by providing meaningful explanations about the RL agent's action policy.

[0016] Figure 1 is a diagram showing an example of an offline RL system 100 according to an embodiment. As shown in Figure 1, the offline RL system 100 has a decision maker 102. During the training phase, the decision maker 102 can receive, as input, past observations from sensors, past actions, and expected future returns, and generate, as output, the next action, a confidence score, and an explanation. After training and during the model application phase, the decision maker 102 can receive the current observation, the most recent action, and the desired action, and generate a prediction based on the input. In some implementation examples, the offline RL system 100 further includes a remaining useful life (RUL) estimator 104. The RUL estimator 104 receives past observations from sensors as input and generates an estimated value of the RUL as input to the decision maker 102.

[0017] The decision maker 102 is a machine learning (ML) model such as a multi-layer perceptron, a convolutional neural network, a transducer, a support vector machine (SVM), a hidden Markov model, a Gaussian process, logistic regression, a gated recurrent unit (GUR), long short-term memory (LSTM), etc., but is not limited thereto. In principle, any ML can perform the work of the decision maker 102. Considering that the input to the decision maker 102 is time series data, ML models adjusted for sequential data such as Transformer and Recurrent Neural Networks (e.g., LSTM, GRU) may exhibit better performance than others. The input to the decision maker 102 consists of the following. For a given window T, at each given time k:

[0018]

Number

[0019]

Number

[0020] By increasing T, more information can be input to the decision maker 102, leading to performance improvement. However, the increase in the feature dimension requires a wider neural network, more hidden layers, a longer training time, etc., making supervised learning computationally more expensive. Future cumulative reward with respect to the horizontal line H (estimated using offline data):

[0021]

Number

[0022]

Number

[0023] For the model that does not use the RUL estimator 104:

[0024]

Number

[0025] For the model that uses the RUL estimator 104 for additional input:

[0026]

Number

[0027] The confidence score c of the model k,It can be generated using at least one of a Bayesian neural network, deep ensemble, Monte Carlo dropout, quantile neural network, etc., but is not limited thereto. The reliability score of the model provides the operator with information regarding the reliability of the RL model and helps determine whether retraining of the RL model is necessary. When the action space is discrete, standard ML classifiers such as support vector machines, QDA, and decision trees provide the probability of each action without additional computational cost. FIG. 8 is a diagram showing an example of the relationship between time and the reliability score of the model. Taking the top diagram of FIG. 8 as an example, if the model generates a reliability score for each action and the recommended action is "do nothing", the generated reliability score is shown to gradually decrease as time progresses. From such a progression, it can be inferred that repair will soon be recommended. Taking the bottom diagram of FIG. 8 as an example, the reliability scores of the model are approximately the same for all actions. Specifically, the model cannot prioritize actions. This may indicate that the operation of the device or machine is new as seen from past records, and promotes the relearning of the machine using newly obtained data.

[0028] The decision maker 102 is trained using past maintenance records and does not require an online interaction with the simulator. Further, the decision maker 102 directly leads to the maintenance decision itself, rather than an indirect signature such as a failure probability or RUL. The reward r is pre-designed by a human operator and reflects the economic costs of repair, replacement, and failure.

[0029] FIG. 2 is a diagram illustrating a predictive maintenance support system 200 according to an embodiment. As shown in FIG. 2, the predictive maintenance support system 200 includes a decision maker 102, a reward calculator 208, a decision maker training engine 210, a database 212, an XAI unit 214, and a graphical user interface (GUI) 218. In some exemplary embodiments, the predictive maintenance support system 200 further includes an RUL estimator 104 and an RUL estimator training engine 216. The RUL estimator 104 is optional, but if the prediction of the RUL estimator 104 is accurate, this will in turn improve the performance of the decision maker 102.

[0030] When receiving instructions for operation or repair from the user 202, the device 204 performs the operation / repair as instructed. Sensors such as internal sensors of the device 204, sensors connected to the device 204, and external sensors of the device 204 generate sensor data by monitoring the performance of the device 204, but are not limited thereto. In some exemplary implementations, the sensor data is received by the sensor data preprocessing unit 206 and performs data processing such as noise removal and dimensionality reduction, but is not limited thereto.

[0031] The processed data generated by the sensor data preprocessing unit 206 is then transferred to the reward calculator 208, the decision maker 102, the RUL estimator 104, and the GUI 218. The sensor data and the processed data may be stored in the database 212. The database 212 stores historical records / data related to past observations, past actions, related rewards, and related predicted actions. The reward generated by the reward calculator 208 may also be stored in the database 212.

[0032] The decision maker training engine 210 and the RUL estimator training engine 216 acquire past records / data from the database 212 and generate the trained decision maker 102 and the RUL estimator 104. The RUL estimator 104 receives the processed data as input and generates an estimated RUL as input to the decision maker 102. The decision maker 102 receives an explanation, the processed data, and the reward associated with the processed data as input and generates the next action and a confidence score. The explanation is generated by the XAI unit 214. The XAI unit 214 accesses both the database 212 and the decision maker 102, analyzes the inside of the decision maker 102, and thereby creates an explanation that can be read by humans. Thereafter, the explanation is transmitted from the decision maker 102 to the GUI 218.

[0033] The XAI unit 214 can provide two types of explanations for the AI's decision. The first is the feature-by-feature explanation, which decomposes the final result of the RL model into the contributions of individual features and identifies the features most responsible for the result. The feature-by-feature explanation can be generated from explanation visualization models such as local interpretable model-agnostic explanations (LIME), Shapley additive explanations (SHAP), and gradient-weighted class activation mapping (Grad-CAM), but is not limited to these. Figure 9 shows two different application examples of the feature-by-feature explanation. As shown in Figure 9, the left figure uses the Grad-CAM method to identify the feature "dog", and the right figure shows the identification of important features in time-series data.

[0034] The second type of explanation is the instance-level explanation, which shows which labeled samples in the training dataset have the most influence on the output of the RL model for a specific input. Figure 10 shows an application example of the instance-level explanation. As shown in Figure 10, when x i is used as the input, y i is generated as the model output. This is the input x iThis is because it is similar to specific identified training samples, such as samples 2, 5, and 6. The explanation at the instance level may be generated from methods such as the influence function, but is not limited to this. The XAI unit 214 greatly facilitates the root cause analysis of equipment failures and shortens the downtime for repairs. For example, when a message such as "We recommend repair. This decision is mainly based on the signal of sensor 2 in the past 10 minutes." is generated, the engineer can know which part of the device detected what kind of abnormality, enabling a more rapid and smooth repair response. Furthermore, when the uncertainty score presented by the AI is high, the engineer can know that additional training of the decision maker 102 using new data is urgently required. The outputs from the decision maker 102 and the sensor data preprocessing unit 206 are received at the GUI 218 and presented to the user 202 there.

[0035] Figure 7 shows an exemplary data table stored in the database 212 according to an exemplary embodiment. As shown in Figure 7, such a data table can include, but is not limited to, information such as sensor data, the revenue associated with each machine, the repair cost associated with each machine, and rewards. The top table identifies the operating states associated with the machine-sensors, and each entry is associated with a unique timestamp. For example, the first entry has a timestamp of "2021 / 12 / 10 13:30", and the state associated with the machine / sensor operating at that time is tracked. For instance, at the timestamp "2021 / 12 / 10 13:30", it indicates that the sensor data of machine ID1 is (12, 0.3, 108.1). The bottom left table identifies the machines / sensors with associated revenues and repair costs, and each entry is associated with a unique timestamp. For example, at the timestamp "2021 / 12 / 10 13:30", the revenue of machine ID1 is "1" and the repair cost is "0". The bottom right table identifies the total reward associated with the timestamp. For example, at the timestamp "2021 / 12 / 10 13:30", it indicates that the total reward is "3.7".

[0036] Figure 3 is a diagram showing an example of the processing flow of the model learning phase of the RL model according to the embodiment. In S302, observation time series data including normal data, failure data, and treatment / repair data is prepared and received. In S304, the data is split into episodes. Rewards are calculated for all time steps and stored in the database 212. In S306, a supervised learning using the observation {o t}, and the true RUL label {RUL t} is used to train the RUL estimator.

[0037] In S308, the decision maker training engine 210 is initialized to train the decision maker 102. A random mini-batch / batch B of the sequence {O, A, R, a} B is sampled by the decision maker training engine 210 in S310, where O, A, R have length T and a is the action taken at the next time step. O is the current observation, A is the action executed in the environment, and R is the associated reward.

[0038] In S312, a loss is calculated by a function:

[0039]

Equation

[0040] In S314, the parameters of the decision maker 102 are updated by gradient descent of the loss to minimize the loss function. In S316, a determination is made as to whether the training of the model has been sufficiently performed. If the answer is no, the process returns to S310 for further training of the model. If the answer is yes, the process ends.

[0041] Figure 4 shows an example of the process flow of the application of the RL model according to the embodiment. At S402, the desired return rate R is received from the user. The desired return-to-go (return rate) is input to the decision maker 102 to distinguish good actions from bad actions. Past records usually contain a mixture of good episodes (low cost) and bad episodes (high cost; too many failures or repairs). By specifying a high return, the RL agent will output actions that closely follow the "best practices" of the dataset. Selecting an unrealistically large value will cause a failure due to uncontrollable extrapolation.

[0042] At S404, new sensor data o and reward r are received from the environment. At S406, a determination is made as to whether the device has failed. If the answer is no, the process proceeds to S410. If the answer is "yes", the process proceeds to S408, where the device is repaired or replaced, and the process proceeds to S410. At S410, R is replaced by R - r.

[0043] At S412, the sensor data / recent observations {o} are supplied to the RUL estimator 104, and the output RUL prediction is received. At S414, the recent sensor data / observations {o}, the recent action {a}, the desired return rate R, and the RUL prediction are input to the decision maker 102. Further, the output from the decision maker 102 including the next action a, the confidence score ck, and the explanation Ek is received. At S416, the output is sent to the GUI 218. At S418, the next action is executed in the environment. At S420, a determination is made as to whether to continue the operation. If the answer is "yes", the process returns to S404 for further processing. If the answer is no, the process ends.

[0044] The foregoing embodiments are considered to have various advantages and merits. For example, a data-driven approach enables the automation of optimized predictive maintenance decisions without requiring domain knowledge of experts. The inference time of a trained RL agent is much shorter than that of conventional mathematical optimization methods. At the same time, since the RL method can operate offline, an expensive high-fidelity equipment simulator is not required. In contrast to methods based on fault likelihood and RUL estimation that require the operator to interpret the output of ML, the offline RL method yields the optimal decision itself. By providing past operation data, complex maintenance scheduling can be performed, such as asynchronous repair of multiple components and state-dependent repair costs, without explicitly modeling individual interdependencies.

[0045] Figure 5 shows an exemplary computing environment having an exemplary computer device suitable for use in some exemplary implementations. The computing device 505 of the computing environment 500 can include one or more processing units, cores, or processors (plural possible) 510, memory 515 (e.g., RAM, ROM, and / or the like), internal storage 520 (e.g., magnetic, optical, solid-state storage, and / or organic), and / or an IO interface 525, any of which can be coupled on a communication mechanism or bus 530 for communicating information or can be embedded in the computing device 505. The IO interface 525 can also be configured to receive images from a camera or provide images to a projector or display, depending on the desired implementation.

[0046] Computing device 505 may be communicatively coupled to an input / user interface 535 and an output device / interface 540. Either or both of the input / user interface 535 and the output device / interface 540 can be a wired or wireless interface and can be removable. The input / user interface 535 can include any device, component, sensor, or interface (physical or virtual) that can be used to provide an input (e.g., buttons, touch screen interfaces, keyboards, pointing / cursor control, microphones, cameras, braille, motion sensors, accelerometers, optical readers, and / or the like). The output device / interface 540 can include a display, television, monitor, printer, speaker, braille, etc. In some exemplary implementations, the input / user interface 535 and the output device / interface 540 can be embedded in the computing device 505 or physically coupled to the computing device 505. In other exemplary implementations, other computer devices can function as or provide the functionality of the input / user interface 535 and the output device / interface 540 of the computing device 505.

[0047] Examples of the computing device 505 include, but are not limited to, highly portable devices (e.g., smartphones, devices mounted on vehicles and other machines, devices carried by humans and animals, etc.), mobile devices (e.g., tablets, notebooks, laptops, personal computers, portable televisions, radios, etc.), and devices not designed for mobility (e.g., desktop computers, other computers, information kiosks, televisions, radios, etc. having one or more processors embedded in and / or coupled thereto).

[0048] Computing device 505 can be communicatively coupled to external storage 545 and network 550 (e.g., via IO interface 525) to communicate with any number of network-connected components, devices, and systems, including one or more computer devices of the same or different configurations. The computing device 505 or any connected computer device can function as a server, client, syn server, general-purpose machine, special-purpose machine, or other label, provide services, or be referenced.

[0049] IO interface 525 can include a wired and / or wireless interface that uses any communication or IO protocol or standard (e.g., Ethernet, 802.11x, Universal System Bus, WiMax, modem, cellular network protocol, etc.) for communicating information between at least all connected components, devices, and networks within computing environment 500, but is not limited thereto. Network 550 can be any network or combination of networks (e.g., the Internet, local area network, wide area network, telephone network, cellular network, satellite network, etc.).

[0050] Computing device 505 can use and / or communicate with computer-usable media or computer-readable media, including transient media and non-transient media. Transient media includes transmission media (e.g., metal cables, optical fibers), signals, carrier waves, etc. Non-transient media includes magnetic media (disks, tapes, etc.), optical media (CD ROM, digital video disks, Blu-ray disks, etc.), solid media (RAM, ROM, flash memory, solid state storage, etc.), and other non-volatile storage or memory.

[0051] Computing device 505 can be used to implement techniques, methods, applications, processes, or computer-executable instructions in some exemplary computing environments. The computer-executable instructions can be obtained from a transient medium, stored in a non-transient medium, and can be obtained from the non-transient medium. The executable instructions can emanate from one or more of programming languages, scripting languages, and machine languages (e.g., C, C++, C#, Java, Visual Basic, Python, Perl, JavaScript, etc.).

[0052] Processor 510 can execute under any operating system (OS) (not shown) in a native environment or a virtual environment. One or more applications can be deployed, including logical unit 560, application programming interface (API) unit 565, input unit 570, output unit 575, and an inter-unit communication mechanism 595 for the different units to communicate with each other, with the OS, and with other applications (not shown). The units and elements described can vary in design, function, configuration, or implementation and are not limited to the provided description. The processor(s) 510 can be in the form of a hardware processor such as a central processing unit (CPU), or a combination of a hardware unit and a software unit.

[0053] In some exemplary implementations, when information or execution instructions are received by the API unit 565, it may be transmitted to one or more other units (e.g., the logic unit 560, the input unit 570, the output unit 575). In some embodiments, the logic unit 560 may be configured to control the flow of information between units and direct the services provided by the API unit 565, the input unit 570, and the output unit 575, as in some of the exemplary embodiments described above. For example, the flow of one or more processes or implementations may be controlled by the logic unit 560 alone or in cooperation with the API unit 565. The input unit 570 may be configured to obtain inputs for the calculations described in the exemplary embodiments, and the output unit 575 may be configured to provide outputs based on the calculations described in the exemplary embodiments.

[0054] The processor(s) 510 can be configured to receive a predicted future return value as an input to the decision maker model, where the decision maker model is a machine learning model that predicts maintenance actions associated with the device, as shown in FIGS. 1-2 and 4. The processor(s) 510 may also be configured to supply recent observations and recent actions from the environment as inputs to the decision maker model, as shown in FIGS. 1-2 and 4. The processor(s) 510 may also be configured to generate the next action as the model output of the decision maker model, where the next action is the predicted maintenance action, as shown in FIGS. 1-2 and 4. The processor(s) 510 may also be configured to execute the next action in an environment such as that shown in FIGS. 1-2 and 4.

[0055] The processor(s) 510 may also be configured to compare the reliability score to a threshold, as shown in FIGS. 1-2. The processor(s) 510 may also be configured to re-learn the machine learning model with the observations that were most recently observed in time compared to the most recent observations and the actions that were most recently observed in time compared to the most recent actions as inputs when the reliability score is below the threshold, as shown in FIGS. 1-2. The processor(s) 510 may also be configured to display the model output on a graphical user interface (GUI), as shown in FIGS. 2 and 4.

[0056] The processor(s) 510 may also be configured to supply the most recent observations as an input to a remaining useful life (RUL) estimator as shown in FIGS. 1-4. The processor(s) 510 may also be configured to generate an estimated remaining useful life of the device as an output from the RUL estimator, as shown in FIGS. 1-4. The processor(s) 510 may also be configured to supply the generated estimated remaining useful life of the device as an input to a decision maker model when generating the next action, as shown in FIGS. 1-2 and 4.

[0057] The processor(s) 510 may also be configured to display the model output and the estimated remaining useful life of the device on a graphical user interface (GUI), as shown in FIGS. 2 and 4. The processor(s) 510 may also be configured to identify a subset of the inputs related to the generation of the model output of the decision maker model, and the subset of the inputs directly affects the generation of the next action, as shown in FIG. 3.

[0058] The processor(s) 510 may also be configured to store data from multiple sensors in a database as the most recent observations and the most recent actions, as shown in FIGS. 1-2. The processor(s) 510 may also be configured to retrieve the most recent observations and the most recent actions from the database, as shown in FIGS. 1-2 and 4.

[0059] Some portions of the detailed description are presented from the perspective of symbolic representations of algorithms and operations within a computer. These algorithmic descriptions and symbolic representations are the means used by those skilled in the data processing arts to convey the essence of their technological innovation to others skilled in the art. An algorithm is a defined sequence of steps leading to a desired final state or result. In an embodiment, the steps executed require physical manipulation of physical quantities to achieve a visible result.

[0060] Unless otherwise specified, as will be apparent from the discussion, throughout this specification, discussions using terms such as "processing," "computing," "calculating," "determining," "displaying," etc., may include operations and transformations of data represented as physical (electronic) quantities within the registers and memories of a computer system to other data similarly represented as physical quantities within the memories or registers of the computer system or other information storage, transmission, or display devices.

[0061] The exemplary embodiments are also related to an apparatus for performing the operations herein. This apparatus may be specially configured for the required purposes, or may include one or more general-purpose computers selectively activated or reconfigured by one or more computer programs. Such computer programs can be stored on a computer-readable medium such as a computer-readable storage medium or a computer-readable signal medium. Computer-readable storage media include tangible media such as optical disks, magnetic disks, read-only memory, random access memory, solid-state devices, drives, or other types of tangible or non-transitory media suitable for storing electronic information, but are not limited thereto. Computer-readable signal media can include media such as carrier waves. The algorithms and displays presented herein are not inherently related to a particular computer or other device. A computer program can include a pure software implementation that includes instructions to perform the operations of the desired implementation.

[0062] Various general-purpose systems may be used with the programs and modules according to the examples herein, or it may prove convenient to construct more specialized devices for performing the desired method steps. Furthermore, the examples are not described with reference to a particular programming language. It will be understood that various programming languages may be used to implement the teachings of the examples described herein. The instructions of the programming language may be executed by one or more processing devices, such as a central processing unit (CPU), a processor, or a controller.

[0063] As is known in the art, the operations described above can be performed by hardware, software, or some combination of software and hardware. Various aspects of the exemplary implementations may be implemented using circuits and logic devices (hardware), while other aspects may be implemented using instructions stored on a machine-readable medium (software) that, when executed by a processor, implement the methods of the present application for the processor. Further, some exemplary implementations of the present application may be performed by hardware only, while other exemplary implementations may be performed by software only. Further, the various functions described may be performed by a single unit or may span multiple components in any number of ways. When performed by software, the methods may be performed by a processor, such as a general-purpose computer, based on instructions stored on a computer-readable medium. Optionally, the instructions may be stored on the medium in a compressed and / or encrypted format.

[0064] Furthermore, other embodiments of the present application will be apparent to those skilled in the art from consideration of the specification and practice of the teachings of the present application. The various aspects and / or components of the described exemplary embodiments may be used alone or in any combination. The specification and exemplary embodiments are intended to be considered as examples only, and the true scope and spirit of the present application are indicated by the following claims.

Claims

1. An offline reinforcement learning method for a predictive maintenance support system for predictive maintenance of equipment, wherein the predictive maintenance support system comprises a database for storing histories related to past observations, past actions, related rewards, and related predicted actions, an XAI unit, a remaining useful life (RUL) estimator, a decision maker training engine, a GUI, a reward calculator, and a decision maker, wherein sensor data of internal sensors of a plurality of devices operating according to instructions from a user is input by the reward calculator, and an expected future revenue expectation value is output, wherein the decision maker inputs the past observation results as the sensor data, the past actions, and the expected future revenue expectation value, and outputs a predicted next action and a confidence score of the decision maker model of the decision maker, wherein the XAI unit accesses the database and the decision maker, and outputs an explanation of the predicted next action that can be read by humans, wherein by the RUL estimator, the past observation results from the sensor are input, and an estimated value of the remaining useful life is output to the decision maker by the RUL estimator, wherein the database stores, for each time, the sensor data, revenue, repair cost, and reward of each of the plurality of devices, wherein the decision maker training engine acquires the data in the database and trains the RUL estimator, wherein further by the decision maker training engine, the confidence score is compared with a threshold value, and when the confidence score is below the threshold value, the decision maker model is retrained with observations newly observed temporally later than the most recent observation and actions newly observed temporally later than the most recent action as inputs An offline reinforcement learning method.

2. The offline reinforcement learning method according to claim 1, wherein the most recent observation value is supplied as an input to the RUL estimator, and an estimated remaining useful life of the equipment is generated as an output from the RUL estimator, and the generated estimated remaining useful life of the equipment is used as an input to the decision maker model of the decision maker when generating the next action An offline reinforcement learning method.

3. The offline reinforcement learning method according to claim 2, wherein the generated estimated remaining useful life of the equipment is displayed on the GUI An offline reinforcement learning method.

4. A non-transitory computer-readable medium storing instructions for predictive maintenance of a machine, wherein the instructions cause a computer to perform the following processes: Input sensor data of internal sensors of a plurality of devices operating according to instructions from a user, and output an expected future revenue expectation value; Input the past observation results as the sensor data, the past actions, and the expected future revenue expectation value; Output a predicted next action and a reliability score of a decision maker model of the computer's decision maker; Access a database storing histories related to past observations, past actions, related rewards, and related predicted actions and the decision maker, and output an explanation of the predicted next action readable by humans; Input past observation results from the sensor and output an estimated remaining useful life value to the decision maker; Store, for each of the plurality of devices at each time, sensor data, revenue, repair cost, and reward; Acquire data from a database storing histories related to past observations, past actions, related rewards, and related predicted actions, and train a remaining useful life (RUL) estimator; Furthermore, Compare the reliability score with a threshold value; When the reliability score is below the threshold value, retrain the decision maker model with observations observed more recently in time than the most recent observation and actions observed more recently in time than the most recent action as inputs; A computer-readable medium. **Claim 5** The computer-readable medium according to claim 4, wherein the instructions Supply a most recent observation value as an input to a remaining useful life estimator (RUL estimator); Generate an estimated remaining useful life of a device as an output from the RUL estimator; The generated estimated remaining useful life of the device is used as an input to the decision maker model when generating the next action; A computer-readable medium. **Claim 6** The computer-readable medium according to claim 5, wherein the instructions Display the model output and the estimated remaining useful life of the device on a graphical user interface (GUI); A computer-readable medium.

Citation Information

Patent Citations

  • Automatic health indicator learning using reinforcement learning for predictive maintenance

    US20190384257A1

  • System for predicting equipment failure events and optimizing manufacturing operations

    US20200265331A1

  • Learning framework for robotic paint repair

    US20210323167A1

  • Approach to determining a remaining useful life of a system

    US20220004182A1

  • Restricting use of selected input in recovery from system failures

    US20220188181A1