Recommendation Background for Operation and Asset Failure Prevention
The system addresses the challenge of predicting failures in industrial systems by using physical and virtual sensor data to generate failure prediction models. This approach enables timely intervention and reduces costs associated with failures, while also identifying root causes automatically.
Patent Information
- Application Number
- JP2024557637
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-03-30
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2042-03-30
AI Technical Summary
Existing industrial systems face challenges in predicting failures, especially long-term ones, due to the rarity of failure events and the reliance on manual processes that are time-consuming and prone to errors. Additionally, the cost and accuracy of collecting fault data pose significant obstacles.
The proposed system collects physical sensor data and generates virtual sensor data using a physics-based model. It identifies key features through sampling, aggregation, and feature derivation operations, and uses machine training models to generate failure prediction models. These models calculate likelihood scores for predicted failures, allowing for timely intervention.
The system effectively predicts both short-term and long-term failures, reducing unplanned downtime and operational delays. It also identifies root causes automatically, facilitating diagnosis and prevention, thereby minimizing financial, reputational, and human costs associated with failures.
Smart Images

Figure 2025514626000001_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure is generally directed to automated fault prediction in industrial systems. [Background technology]
[0002] Automated fault prediction may be relevant to many complex industrial systems in many different industrial contexts (e.g., industries). For example, different industrial contexts may include, but are not limited to, manufacturing, entertainment (e.g., theme parks), hospitals, airports, utilities, mining, oil and gas, warehouses, and transportation systems. Two main fault types may be defined by how far away a fault is in terms of time from its symptom. A short-term fault type may relate to a symptom and a fault that is close in time (e.g., hours or days). For example, an overload fault on a conveyor belt may be a short-term fault type associated with a symptom that is close in time. A long-term fault type may relate to a symptom that is distant in time (e.g., weeks, months, or years) from the fault. Long-term fault types of system component failures often progress gradually and chronically and may have broader negative impacts, such as a shutdown of the entire system including the faulty component. For example, a long-term fault may include a component failure due to breaks and cracks in a dam or metal fatigue.
[0003] Although failures in industrial systems may be rare, the costs of such failures may result in significant (or significant) financial (e.g., operations, maintenance, repairs, logistics, etc.), reputational (e.g., marketing, market share, sales, quality, etc.), human (e.g., scheduling, skill set, etc.), and / or liability (e.g., safety, health, etc.) costs. Some industrial systems may not have a fault detection process in place. In other industrial systems, fault detection is performed manually based on domain knowledge based on real-time sensor data. Manual fault prediction and / or detection may pose challenges when dealing with high frequency data (e.g., vibration sensor data, IoT data, etc.).
[0004] Even if a failure can be detected and / or predicted, it may be too late to repair or recover, since the failure may have already occurred or be very close to occurring. Therefore, there is a need to predict failures at a time when an operator or technician may have enough time to respond to, repair, or even avoid the failure. Proactive failure prediction can facilitate the avoidance of failures and the reduction of losses and / or negative impacts resulting from the failure.
[0005] Fault data, in some aspects, is expensive to collect, and data on faults may not be collected at all, or the fault data may be inaccurate, incomplete, and / or unreliable. This poses challenges to building supervised solutions that rely on fault data as labels. Thus, hereinafter, a system is presented that collects appropriate data and uses the collected data to predict one or more of short-term or long-term faults in time to reduce or avoid costs associated with the predicted faults. In some aspects, the system may further identify a set of root causes of the faults, and may use the root causes as further information to facilitate diagnosis of the faults and take steps to repair or avoid the faults. In some aspects, the root cause analysis may be performed by the system without manual inspection of the components or visualization of the sensor data and metrics (e.g., because manual inspection may be time-consuming, expensive, and / or error-prone). The set of root causes, in some aspects, is a set of root causes identified at the sensor data level (e.g., data collected on temperature, vibration levels, or some other measured / monitored characteristic associated with the fault). Summary of the Invention [Means for solving the problem]
[0006] Exemplary implementations described herein include an innovative method. The method may include collecting a physical sensor data set. The method may further include generating a virtual sensor data set by applying a physics-based model to a subset of the physical sensor data set. The method may also include identifying a first set of features from the physical sensor data set and the virtual sensor data set by performing at least one of a sampling operation, an aggregation operation, or a feature derivation operation on the physical sensor data set and the virtual sensor data set based on an optimized sampling rate or an optimized aggregation statistic. The method may further include identifying a set of anomaly detection scores, a first set of contributors to the set of anomaly detection scores, and a second set of features having a feature importance score that is equal to or greater than a threshold using a first set of machine training models applied to the first set of features. The method may further include generating at least one failure prediction model based on the first set of machine training models. The method may also include applying the at least one failure prediction model to the physical sensor data set and the virtual sensor data set to calculate a likelihood score for a predicted failure of the asset.
[0007] An exemplary implementation described herein includes an innovative computer-readable medium storing computer-executable code. The computer-executable code may include instructions for collecting a physical sensor data set. The computer-executable code may also include instructions for generating a virtual sensor data set by applying a physics-based model to a subset of the physical sensor data set. The computer-executable code may further include instructions for identifying a first set of features from the physical sensor data set and the virtual sensor data set by performing at least one of a sampling operation, an aggregation operation, or a feature derivation operation on the physical sensor data set and the virtual sensor data set based on an optimized sampling rate or an optimized set of aggregation statistics. The computer-executable code may also include instructions for identifying a set of anomaly-detection scores, a first set of contributors to the set of anomaly-detection scores, and a second set of features having feature importance scores that are equal to or greater than a threshold value using a first set of machine training models applied to the first set of features. The computer-executable code may also include instructions for generating at least one fault prediction model based on the first set of machine training models. The computer executable code may further include instructions for applying at least one failure prediction model to the set of physical sensor data and the set of virtual sensor data to calculate a likelihood score for a predicted failure of the asset.
[0008] Exemplary implementations described herein include an innovative apparatus. The apparatus may include a memory and at least one processor configured to collect a physical sensor data set. The at least one processor may also be configured to generate a virtual sensor data set by applying a physics-based model to a subset of the physical sensor data set. The at least one processor may further be configured to identify a first set of features from the physical sensor data set and the virtual sensor data set by performing at least one of a sampling operation, an aggregation operation, or a feature derivation operation on the physical sensor data set and the virtual sensor data set based on an optimized sampling rate or an optimized aggregation statistic. The at least one processor may also be configured to identify a set of anomaly detection scores, a first set of contributors to the set of anomaly detection scores, and a second set of features having feature importance scores that are equal to or greater than a threshold using a first set of machine training models applied to the first set of features. The at least one processor may also be configured to generate at least one fault prediction model based on the first set of machine training models. The at least one processor may be further configured to apply at least one failure prediction model to the set of physical sensor data and the set of virtual sensor data to calculate a likelihood score for a predicted failure of the asset. [Brief description of the drawings]
[0009] [Figure 1] FIG. 1 is a diagram illustrating the basic concept of a circular economy according to some embodiments of the present disclosure.
[0010] [Figure 2A] FIG. 2A is a diagram illustrating components of a physics-based model according to some aspects of the disclosure.
[0011] [Figure 2B] FIG. 2B illustrates a feature engineering module according to some aspects of the disclosure.
[0012] [Diagram 3] FIG. 3 is a flow diagram illustrating an example workflow for optimizing at least one of a sampling rate and an aggregate statistic according to some aspects of the present disclosure.
[0013] [Figure 4] FIG. 4 illustrates a set of components of an implementation of an anomaly detection module (eg, an unsupervised anomaly detection model, etc.) according to some aspects of the disclosure.
[0014] [Diagram 5] FIG. 5 illustrates a set of operations related to short-term fault prediction.
[0015] [Figure 6] FIG. 6 illustrates a set of operations associated with generating a long-term disability prediction score for identifying and predicting potential long-term disability, according to some embodiments of the present disclosure.
[0016] [Figure 7] FIG. 7 illustrates a set of operations for identifying a set of root causes associated with short-term disorders in accordance with some aspects of the present disclosure.
[0017] [Figure 8] FIG. 8 is a flow diagram illustrating a set of operations for identifying a set of contributing factors associated with long-term impairment for assets in one or more asset classes, according to some embodiments of the present disclosure.
[0018] [Figure 9] FIG. 9 is a flow diagram illustrating a system for performing fault prediction operations for an industrial process.
[0019] [Figure 10] FIG. 10 is a flow diagram illustrating a method for identifying a set of root causes (eg, contributing factors) for failure prediction according to some aspects of the present disclosure.
[0020] [Figure 11] FIG. 11 illustrates an example computing environment having an example computer device suitable for use in some example implementations. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0021] The following detailed description shows the drawings and implementation details of the present application. Reference numbers and descriptions of redundant elements between the drawings are omitted for clarity. The terms used throughout the description are provided as examples and are not intended to be limiting. For example, the use of the term "automatic" may include a fully automatic implementation or a semi-automatic implementation that includes user or administrator control over certain aspects of the implementation, depending on the desired implementation of the person skilled in the art practicing the implementation of the present application. The selection may be made by a user through a user interface or other input means, or may be implemented by a desired algorithm. The implementations described herein may be utilized alone or in combination, and the functionality of the implementations may be implemented by any means depending on the desired implementation.
[0022] In this disclosure, systems, apparatus, and methods are presented that address the problem of automated prediction of faults (short-term and long-term faults) and identification of the root causes of predicted faults in industrial systems with unlabeled high-frequency sensor data.
[0023] In the present disclosure, systems, apparatuses, and methods are presented that provide techniques for fault prediction operations for industrial processes. For example, the method may include collecting a physical sensor data set. The method may further include generating a virtual sensor data set by applying a physics-based model to a subset of the physical sensor data set. The method may also include identifying a first set of features from the physical sensor data set and the virtual sensor data set by performing at least one of a sampling operation, an aggregation operation, or a feature derivation operation on the physical sensor data set and the virtual sensor data set based on an optimized sampling rate or an optimized aggregation statistic. The method may further include identifying a set of anomaly detection scores, a first set of contributors (e.g., root causes) to the set of anomaly detection scores, and a second set of features having feature importance scores that are equal to or greater than a threshold value using a first set of machine training models applied to the first set of features. The method may also include generating at least one fault prediction model based on the first set of machine training models. The method may further include applying at least one failure prediction model to the set of physical sensor data and the set of virtual sensor data to calculate a likelihood score for a predicted failure of the asset.
[0024] In some aspects, the systems, apparatus, and methods described herein may be directed to prediction of both short-term and long-term failures and derivation of root causes of predicted failures to mitigate or avoid negative impacts prior to failure. There may be several key benefits of failure prediction and prevention solutions. Failure prediction and prevention solutions, in some aspects, may reduce unplanned downtime and operational delays while increasing productivity, output, and operational efficiency. In some aspects, failure prediction and prevention solutions may optimize yields and increase margins / profits.
[0025] Fault prediction and prevention solutions, in some aspects, may maintain production consistency and product quality. In some aspects, fault prediction and prevention solutions may reduce unplanned costs for logistics, maintenance scheduling, labor, and repair expenses. Fault prediction and prevention solutions, in some aspects, may reduce damage to assets and the entire industrial system. In some aspects, fault prediction and prevention solutions may reduce accidents to operators and improve operator health and safety. The proposed solutions generally provide benefits to all entities involved in the industrial system, including, but not limited to, operators, supervisors / managers, maintenance technicians, SMEs / domain experts, assets, and the system itself.
[0026] Some problems (limitations and constraints) of conventional systems and methods are described below. Techniques to solve these problems are described herein. For example, conventional systems / methods may be highly dependent on accurate historical failure data. However, accurate historical failure data is usually not available for several reasons. For example, historical failure-related data may not be collected or may be inaccurate or incomplete, processes for collecting failure data may not exist or may not collect enough data to perform useful analysis, and / or collected data (e.g., IoT data) may be too voluminous for manual processing, detection, and identification of failure data.
[0027] Additionally, in some aspects, there may be no standard process for effectively and efficiently detecting and classifying both common and rare events. The manual process of collecting faults by labeling sensor data based on domain knowledge may be inaccurate, inconsistent, unreliable, and time-consuming in some aspects. The process in place to collect fault-related data by an operator of the industrial system may not be complete enough in some aspects to identify and investigate root causes. In some aspects, insufficient data collection may often be due to a lack of understanding at the time of data collection as to how the data may help identify root causes. Thus, some aspects provide an automated standard process or method for accurately, effectively, and efficiently detecting and collecting faults in an industrial system.
[0028] Conventional failure prediction solutions, in some aspects, do not perform well in the case of rare failure events with associated lead times (e.g., sufficient lead time to implement a solution to repair or prevent the predicted failure). In some aspects, failure prediction solutions may not perform well for rare events because sensor data may be highly biased toward good operating conditions. Due to the rarity of such failures, in some aspects, it is very difficult to build supervised machine learning modeling with high accuracy and sensitivity. For example, in some aspects, it may be difficult to determine one or more optimal windows for collecting features / evidence regarding a failure or identifying signals that can be used to predict a failure.
[0029] In some aspects, there may not be enough data to identify patterns from a limited amount of fault data. For example, industrial systems usually run in normal conditions and faults are usually rare events, and it may be difficult to capture patterns for a limited number of faults and therefore difficult to predict such faults. Thus, in some aspects, it may be difficult to build a correct relationship between normal cases and rare fault events in a time sequence because it is difficult to capture a sequence pattern of the progression of rare faults. Thus, the system, apparatus, and / or method may identify the correct signals (e.g., features) for fault prediction within an optimal feature window to provide the fault prediction with enough lead time to address the predicted fault. The system, apparatus, and / or method may have the ability to build / identify the correct relationship between normal cases and rare faults and the progression of rare faults in one or both of short-term and long-term faults.
[0030] The system, apparatus, and / or method may provide automated root cause analysis. Root cause analysis of a fault may be performed manually based on domain knowledge and data visualization, which may be subjective, time-consuming, and prone to error. In some cases, the root cause may be associated with raw sensor data that is not addressed by the domain knowledge or data visualization used for the manual root cause analysis. The system, apparatus, and / or method may provide automated root cause based on a standardized method for identifying predicted root causes of faults and outputting the root causes at different levels (including the raw sensor data level).
[0031] In some aspects, the sensor data (e.g., IoT sensor data, vibration data) may be high frequency data (e.g., 1000 Hz to 3000 Hz). High frequency data, in some aspects, poses challenges in constructing a solution to the fault prediction problem. For example, high frequency data may be associated with high levels of noise or long or resource consuming analysis (e.g., computation) times. The sampling frequency or aggregation window may require optimization to accurately predict one of the short term or long term faults. Thus, the system, apparatus, and / or method may provide a window optimization operation to identify an optimized window and / or aggregation statistic for fault prediction.
[0032] In some aspects, physical sensor data may not be able to capture all of the signals that may be useful for monitoring the system due to the harsh environment of the sensor installation, the cost of the sensor, and / or the functionality of the sensor. As a result, the collected data may not be sufficient to monitor the health of the system and capture potential risks and failures. In some aspects, the inability to capture all potentially useful signals may pose a challenge to building a failure prediction solution. Thus, the system, apparatus, and / or method may augment the physical sensor data to capture the necessary signals to help the system monitor and build a failure prediction solution. For example, the physical sensor data may be processed by a set of physics-based models to generate virtual sensor data.
[0033] In some aspects, the system, apparatus, and / or method may implement and / or include several techniques to generate one or more fault prediction models and / or identify a set of root causes. The techniques may include semi-empirical methods that utilize one or more of physics-based models and data-driven machine learning models. In some aspects, physics-based models may be used to augment sensor data with additional features based on the physics of the system. For example, torques on system components (e.g., joints on a robotic arm) may be calculated based on a set of physical data sensors including one or more of accelerometers (linear or rotational), force sensors, IoT motors, or other such sensors associated with the associated components (e.g., motors, boom ends, etc.). Additionally, sampling and aggregation optimization methods may be used to sample and aggregate high frequency data to derive features from the aggregated data.
[0034] The techniques may further include one or more unsupervised fault prediction techniques and / or solutions. In some aspects, the unsupervised fault prediction techniques and / or solutions may be based on sensor data without relying on historical fault data. For example, an unsupervised ensemble anomaly detection model may be used to derive (1) anomaly scores as labels and / or (2) features for the fault prediction model. The techniques and / or solutions may further include supervised surrogate models for the anomaly detection model for feature selection, root cause analysis, and model evaluation.
[0035] In some aspects, the short-term failure prediction model may be based on the ensemble anomaly scores, the selected features, and the root causes identified using the anomaly detection model, and may use one or more window-based feature derivation techniques to derive aggregated features and predict failures with lead times, for example by using machine learning, such as a deep learning sequence prediction model (such as a long short-term memory (LSTM) or gated recurrent unit (GRU)). The long-term failure prediction model may in some aspects be based on the aggregated ensemble anomaly scores, the selected features, and the root causes from the anomaly detection model for each asset in the set of assets of the system. As a result, the long-term failure prediction model may be built based on aggregated features from multiple assets.
[0036] In some aspects, root cause analysis of short-term and long-term faults may be further performed to identify root causes of predicted faults (short-term and long-term faults) at different levels. For example, root causes may be identified based on detected anomaly scores, selected feature sets, sensor data (e.g., physical or virtual sensor data) by using a chain of explainable AI models and aggregation / ranking algorithms.
[0037] FIG. 1 is a diagram 100 illustrating conceptual elements of a solution architecture for semi-empirical unsupervised fault prediction and root cause analysis according to some aspects of the disclosure. Diagram 100 includes a set of sensor data 110 (e.g., radio frequency data from IoT sensors) collected from multiple physical sensors. The physical sensor data 110 may be provided to a physics-based model 120. The physics-based model 120 may be applied to the physical sensor data (or a subset of the physical sensor data) to generate virtual sensor data (e.g., data that enhances a dataset for fault prediction and root cause analysis). The physical sensor data and the virtual sensor data may be provided to a feature engineering module 130 to sample and aggregate the sensor data (physical and virtual) and derive new features based on an optimized sampling rate and / or an optimized set of aggregate statistics / characteristics. In some aspects, the feature engineering module 130 may perform an optimization function (e.g., machine learning based optimization) for one or more of the sampling rate or the set of aggregate statistics / characteristics prior to execution. For example, in some aspects, one of the sampling rate or the aggregate statistics / characteristics may be provided by a user based on domain knowledge, while the other of the sampling rate or the aggregate statistics / characteristics may be optimized.
[0038] The data generated and / or processed by the feature engineering module 130 may be provided to the anomaly detection module 140. The anomaly detection module 140 may use the data from the feature engineering module 130 to build multiple anomaly detection models and ensembles of anomaly detection models. The anomaly detection module 140 may build multiple anomaly detection models and ensembles of anomaly detection models using machine learning operations in a first learning phase and use these in a second prediction (inference) phase. The anomaly detection module 140 may use multiple anomaly detection models and ensembles of anomaly detection models to generate an ensemble anomaly score, derive root causes (e.g., associated physical or virtual sensors) for each data point, and select features through surrogate supervised models for the ensemble anomaly detection models.
[0039] The data processed by the anomaly detection module 140 may be provided to a short-term fault prediction module 150. The short-term fault prediction module 150 may derive features by a look-back feature window and may use a deep learning sequence prediction model (e.g., LSTM or GRU) to predict faults in advance. Based on the output of the short-term fault prediction module 150, the system may identify and / or derive a set of root causes using root cause analysis of the short-term fault module 160. For example, for each predicted fault, the root cause analysis of the short-term fault module 160 may derive the root causes by a chain of explainable AI techniques and ranking / aggregation algorithms.
[0040] Similarly, the data processed by the anomaly detection module 140 may be provided to the long-term failure prediction module 170. The long-term failure prediction module 170 may derive features by aggregation techniques based on the anomaly scores, root causes, and selected features. The long-term failure prediction module 170 may use at least one further anomaly detection model to identify and predict long-term failures. Based on the output of the long-term failure prediction module 170, the system may identify and / or derive a set of root causes using the root cause analysis of the long-term failure module 180. The root cause analysis of the long-term failure module 180 may, in some aspects, build a surrogate supervised model for another anomaly detection module. In some aspects, for each predicted failure, the root cause analysis of the long-term failure module 180 may use a surrogate supervised model to derive a set of root causes by a chain of explainable AI techniques and ranking / aggregation algorithms.
[0041] The following sections provide a detailed description of each component in the solution architecture. The specific methodologies used in connection with the feature engineering module 130, the anomaly detection module 140, the short term fault prediction module 150, the root cause analysis of the short term faults module 160, the long term fault prediction module 170, and the root cause analysis of the long term faults module 180 are described below.
[0042] In some aspects, the sensor data 110 may be collected by IoT sensors that are installed on a set of assets of interest and used to collect data to monitor the health status and performance of the assets and the overall system. Different types of sensors are designed to collect different types of data for different industries, different assets, and / or different tasks. In the following description, the sensors are generally described with the assumption that the methods of data processing are applicable to different types of sensor data with minor adjustments. Some examples of sensors that may be used to collect the sensor data 110 may include temperature sensors, pressure sensors, vibration sensors, acoustic sensors, motion sensors, optical sensors, LIDAR sensors, infrared (IR) sensors, acceleration sensors, gas sensors, smoke sensors, humidity sensors, level sensors, image sensors (cameras), proximity sensors, water quality sensors, and / or chemical sensors.
[0043] For a particular target asset, in addition to the sensors installed on the particular target asset, other sensors installed in the system may be used to build models (e.g., anomaly detection models, failure prediction models, and / or root cause models) for the particular target asset. For example, data collected from a set of sensors installed on the asset or on system components upstream of the particular target asset and / or downstream of the particular target asset may be used to build a set of failure prediction models and / or identified as associated with a predicted root cause of failure. The selection of sensors to be considered in building different models may be based on domain knowledge in some aspects. In some aspects, a full set of sensors may be used to generate different models for the particular target asset, and feature detection / selection and root cause analysis may be used to narrow down the set of sensors associated with the set of trained models (e.g., failure prediction models). For example, data analytics and model-based feature selection may be applied to select sensors associated with the failure prediction models.
[0044] 2A is a diagram 200 illustrating components of a physics-based model according to some aspects of the disclosure. The physics-based model 220 corresponds in some aspects to the physics-based model 120 of FIG. 1. As discussed above, a physical sensor may not capture a complete set of relevant signals and / or metrics to assist in monitoring the health of a system. The inability to capture a complete set of relevant signals and / or metrics may be due to one or more reasons. For example, a physical sensor may not capture a set of expected signals due to physical limitations of the hardware, a physical sensor may not be installable in a harsh environment, such as, for example, inside a location that has high levels of radiation or has pressures or temperatures that are outside the range in which the sensor can function.
[0045] In some aspects, the set of physical sensors may not capture data at the expected frequency. To overcome the limitations of the physical sensors, in some aspects, software-based methods may be used to obtain the expected signals. The physics-based model 224 may in some aspects be a representation of laws of nature that inherently incorporate concepts of time, space, causality, and generalizability. These laws of nature define how physical, chemical, biological, and geological processes evolve in some aspects. The physics-based model 224 may in some aspects be a function that takes multiple inputs (e.g., physical sensor data 222) and generates multiple outputs (e.g., virtual sensor data 226). The inputs may come from predefined profiles (such as motion profiles) at design time or from physical sensors (e.g., physical sensor data 222) during operation. The outputs (e.g., virtual sensor data 226) may in some aspects include multiple variables and represent the set of virtual sensors. In some aspects, a virtual sensor is a type of software-derived information from available information that represents or is associated with data that a corresponding physical device would collect. For example, data collected by a physical acceleration sensor may be combined with the known mass of an associated component to calculate the (virtual) output of a virtual force sensor of the associated component. In some aspects, the virtual data associated with a virtual sensor may be used in the same manner as data from a physical sensor to derive insights and / or build models and / or solutions for downstream tasks.
[0046] Benefits of using virtual sensors may include data capture and data validation in some aspects. Data capture may include using virtual sensors to collect data (e.g., capture signals) that cannot be captured by physical sensors. For example, physical sensors may not be able to capture data due to hardware limitations of the physical sensors or harsh environments that are not compatible with the installation or capabilities of the physical sensors. In some aspects, when there is not much data available in the early stages of physical sensor installation, virtual sensors may be used to derive insights and build models. Virtual sensors may generate high frequency data that a set of physical sensors may not be able to capture in some aspects. Data validation may also be used in some aspects to validate physical sensor data when collecting the same or correlated sets of data. For example, virtual sensor data may be used as an "expected" value while physical sensor data may be used as an "observed" value, and discrepancies or differences between them may be used as a signal to detect abnormal behavior or anomalies in the system.
[0047] Through virtual sensors, in some aspects, physics-based models and machine learning models may be combined as a semi-empirical approach, which can take advantage of both domain knowledge (through physics-based models) and data-driven methods (through machine learning models). Physics models are theoretically self-consistent and have demonstrated success in providing experimental predictions. Physics-based models usually work well during system design. However, during operation, complex system interactions and situations are involved, and theoretical physics-based models based on domain knowledge and simulations may be unable to capture the underlying mechanisms and may be more inaccurate and insensitive. Meanwhile, data-driven methods may capture subtle signals and patterns in complex systems (if sufficient data is collected) and derive appropriate insights for decision-making. Physics-based models can complement machine learning models by incorporating domain knowledge into artificial intelligence (AI) and / or machine learning (ML) models that may be expensive to discover based on pure data-driven methods. 2A illustrates how the physics-based model 224 is applied to the physical sensor data 222 in a time series format to derive the virtual sensor data 226. In some aspects, only a subset of the physical sensor data (e.g., a subset of the sensor data 110 of FIG. 1) may be used as input physical sensor data 222 to the physics-based model 224. The physical sensor data 222 may be pre-processed in some aspects to be fed to the physics-based model 224. For example, the physical sensors may capture the position of the asset as it moves, but the physics-based model may be based on the velocity and acceleration of the asset as it moves. Thus, the pre-processing or physics-based model may include a first derivative that is used to calculate velocity data associated with the position data, and may include a second derivative that is used to calculate acceleration data associated with the position data. The data from both the physical and virtual sensors will be used as input to the next module in the solution architecture (e.g., the feature engineering module 130 of FIG. 1).The physics-based model may in some aspects be constructed by simulation software and / or tools and may also be constructed based on domain knowledge.
[0048] 2B is a diagram 250 illustrating a feature engineering module 230 according to some aspects of the disclosure. The feature engineering module 230, in some aspects, is one implementation of the feature engineering module 130 of FIG. 1. In some aspects, some feature engineering techniques may be deployed to derive features from sensor data (e.g., physical or virtual sensor data, radio frequency IoT data, etc.). FIG. 250 illustrates some steps within the feature engineering module 230. Both physical and virtual sensor data may be used to derive features in some aspects.
[0049] Because the sensor data may in some aspects be high frequency (such as 1000 Hz or 3000 Hz) time series data, there may be a downsampling operation to convert the data to a lower frequency and / or an aggregation operation to aggregate the data in order to capture useful signals for downstream tasks and / or analysis and use some techniques to derive features from the low frequency and / or aggregated data. The downsampling and / or aggregation operation may be performed in a sampling and aggregation module 234 that receives the sensor data (e.g., physical and virtual sensor data 232). In some aspects, an output of the sampling and aggregation module 234 may be provided to a feature derivation module 236 to identify a first set of features 238.
[0050] In some aspects, a sampling rate and aggregation statistics / characteristics may be determined prior to performing downsampling and / or aggregation operations for the high frequency sensor data. For example, the sampling rate may relate to the amount of data to be retained when providing the data to the feature derivation module 236. For example, if the sampling rate is 0.01, 1 percent of the original data will be retained in the resulting data, while if the sampling rate is 0.1, 10 percent of the original data will be retained in the resulting data for feature derivation.
[0051] To compensate for data loss due to the downsampling operation, in some aspects, aggregate statistics / characteristics may be provided for the original high frequency sensor data over each of a set of time windows. For example, the aggregate statistics / characteristics may include, but are not limited to, a minimum value in the time window, a maximum value in the time window, an average value in the time window, a standard deviation, a value associated with the 1st percentile, a value associated with the 99th percentile, a value associated with the 25th percentile, a value associated with the 50th percentile, a value associated with the 75th percentile, and a trend. For example, when sampling data between 1000 Hz (1000 data points / second) and 1 Hz (1 data point / second), some statistics will be calculated over the 1000 Hz data every second and such statistics will be provided in the resulting data provided to the feature derivation module 236.
[0052] Sampling rates and aggregate statistics may be suggested in some aspects based on domain knowledge. However, the suggested values may not be optimal for downstream solutions. Therefore, optimization methods for optimizing sampling rates and aggregate statistics for downstream solutions may be performed before any downsampling or aggregation operations for real-time predictive failures (e.g., in the estimation phase after models are trained and root causes are identified).
[0053] 3 is a flow diagram 300 illustrating an example workflow for optimizing at least one of a sampling rate and an aggregate statistic according to some aspects of the disclosure. The optimization may be performed by the feature engineering module 130 or 230 (or more particularly the sampling and aggregation module 234). At 302, a set / list of aggregate statistics may be generated that identifies statistics and / or characteristics included in the results of the aggregation operation. As mentioned above, the aggregate statistics / characteristics may include, but are not limited to, a minimum value in a time window, a maximum value in a time window, an average value in a time window, a standard deviation, a value associated with the 1st percentile, a value associated with the 99th percentile, a value associated with the 25th percentile, a value associated with the 50th percentile, a value associated with the 75th percentile, and a trend. The space of possible aggregate statistics and sampling rates may be explored in some aspects based on at least one of a number of optimization methods (e.g., Bayesian optimization, grid search, or random search).
[0054] At 304, the workflow may include randomly selecting a subset of aggregate statistics from the set / list of aggregate statistics generated at 302. Randomly selecting the subset of aggregate statistics from the set / list of aggregate statistics may include generating a random binary value (e.g., an N-bit binary value) of length N, where N is the number of elements in the set / list of aggregate statistics such that elements of the set / list of aggregate statistics corresponding to a "0" in the N-bit binary value are not included in the randomly selected subset of aggregate statistics, while elements of the set / list of aggregate statistics corresponding to a "1" value are included in the randomly selected subset of aggregate statistics.
[0055] At 306, the workflow may include selecting a sampling rate. The sampling rate may be selected at 306 either randomly or based on one of domain knowledge or business requirements. For example, the sampling rate may be selected based on knowledge of characteristic time scales associated with asset failures.
[0056] At 308, the workflow may include performing sampling and aggregation operations on the sensor data (e.g., physical sensor data and / or virtual sensor data). The subset of aggregate statistics and sampling rates used to perform the sampling and aggregation operations at 308 may in some aspects be a subset of the aggregate statistics selected at 304 and the sampling rate selected at 306. Performing the sampling and aggregation operations may provide data for feature derivation operations performed by feature derivation module 236, as described in connection with feature engineering module 230 of FIG. 2B.
[0057] The workflow may then include, at 310, building at least one model based on the sampled and aggregated sensor data (and / or any identified features in the first set of identified features). The at least one model may include one or more of an anomaly detection / scoring model, a short-term fault prediction model, or a long-term fault prediction. The model built based on the subset of aggregate statistics selected at 304 and the sampling rate selected at 306 may be evaluated based on a number of set model performance metrics, such as overall accuracy, precision, and recall. In some aspects, building the model at 310 may include training a machine learning model based on the sampled and aggregated data.
[0058] At 312, the workflow may include determining whether a set and / or desired number of iterations of sampling rate and aggregate statistic selection has been reached. The set number of iterations may be based on a grid-based search of the sampling rate-aggregate statistic space or some other characteristic of the system or model. If the workflow determines at 312 that the set and / or desired number of iterations has not been reached, the process may return to 304 to select another subset of aggregate statistics.
[0059] However, if the workflow determines that the set and / or desired number of iterations has been met, the workflow may proceed to train a model (e.g., a Gaussian regression model) at 314. The trained model may be a surrogate model based on the results from 310. In some aspects, the features are binary representations of aggregate statistics plus a sampling rate, and the target may relate to a performance metric. The surrogate model may be one of a Gaussian Process model or a Tree Parzen Estimator (TPE) model for Bayesian optimization. In some aspects, if the machine learning model for the downstream task is too complex, a simpler machine model (linear model, tree-based model) may be used as a surrogate to / for the complex machine learning model.
[0060] At 316, the workflow may include defining an acquisition function to aid in selecting an optimized set of features (binary representation of aggregate statistics plus sampling rate) for the Gaussian regression model. In some aspects, the acquisition function may be one of probability of improvement, expected improvement, Bayesian expected loss, upper confidence limit (UCB), Thompson sampling, or a hybrid of one or more of these. In some aspects, each different acquisition function may be associated with a different trade-off between exploration and exploitation and may be selected to minimize the number of feature queries.
[0061] At 318, the workflow may include training one or more machine learning models with the optimal set of values obtained based on the acquisition function employed at 316. The training may include identifying a performance metric and a run time at 318. Based on the output of the training of the one or more components 318, the workflow may determine whether the run time is equal to or greater than a threshold time. If the run time is determined to be equal to or greater than the threshold at 320, in some aspects the workflow may return to randomly select a subset of aggregate statistics from the set / list of aggregate statistics at 304. If the run time is determined to be less than the threshold at 320, the workflow may further determine whether one or more criteria for stopping (e.g., stopping criteria) are met at 322. If the workflow determines that the stopping criteria are not met (e.g., the training performance metrics do not meet a predefined criterion) at 322, the last selected optimal set of features (binary representation of the aggregate statistics plus sampling rate) and the model performance metrics may be added to a training dataset for the Gaussian regression model at 324, and the workflow may return to another round of training at 314 as shown. If the workflow determines that a stopping criterion is met, at 322, the workflow may terminate. In some aspects, the criterion may relate to the number of rounds, a model metric, a variance or entropy reduction rate, or other related criteria.
[0062] In some embodiments, there may be different algorithms used in the optimization process at different stages of the workflow depicted in FIG. 3. For example, while the optimization method described above relates to Bayesian optimization, in some embodiments, other optimization methods such as grid search (rough search) and random search may be used. Similarly, while in general, Gaussian process models may be used as surrogate functions for Bayesian optimization, the workflow may use other surrogate functions for a particular business / industry problem. As described above, in some embodiments, the workflow may utilize simpler machine models (e.g., linear models or tree-based models) as surrogates for complex machine learning models when the machine learning models for the TPE or downstream tasks are too complex. In addition, different acquisition functions may be implemented at 316, including, but not limited to, probability of improvement, expected improvement, Bayesian expected loss or UCB, Thompson sampling, and / or one or more hybrids of acquisition functions. Each different acquisition function may be associated with a tradeoff between exploration and exploitation, and a particular acquisition function may be selected to minimize the number of feature queries.
[0063] In some aspects, two phases of optimization may be defined. The first phase may include determining aggregate statistics (e.g., a selected subset of aggregate statistics) and optimizing a sampling rate, while the second phase may include determining a sampling rate to optimize the set of aggregate statistics. In each phase, the determined parameters (e.g., the sampling rate or one of the subsets of aggregate statistics) may be determined based on domain knowledge. In some aspects, the first and second phases may be performed iteratively or in different orders. For example, based on domain knowledge, a sampling rate (or a set of possible sampling rates) may be identified, and the second phase of optimization may be performed according to the workflow of FIG. 3. For example, in the second phase, at 306, a sampling rate may be selected based on the sampling rate or set of sampling rates identified based on domain knowledge. After optimizing the subset of aggregate statistics, the first phase may be performed (e.g., first or ith) to optimize the sampling rate for the optimized subset of aggregate statistics identified in the second phase. The first and second phases may be iterated to perform a set of "one-dimensional" optimization operations (e.g., only one set of varying parameters) to converge on an optimized set of parameters for sampling rate and aggregate statistics.
[0064] Based on the optimized sampling rate and the aggregated statistics, a set of historical sensor data (e.g., physical and / or virtual sensor data) may be sampled and aggregated to generate processed sensor data. The processed sensor data may be used, for example, by feature engineering module 130 or 230 (or feature derivation module 236) to derive a first set of features by applying techniques to the time series data. The techniques may include, but are not limited to, one or more of moving average, moving variance, differentials associated with the rate of change of values of the time series data (e.g., for the first or second derivative).
[0065] Additionally, for each time point, the feature detector may define a lookback feature window to derive some statistics (e.g., in a set-off aggregate statistics) about the data within the feature window, and may use the derived statistics as further features at the current time point (e.g., associate the derived statistics with the time point of the refined data set). The length of the feature window may be determined (at least initially) based on domain knowledge, and may be further optimized by optimization techniques (such as grid search, random search, or Bayesian optimization, as described above in connection with FIG. 3). Such derived features may in some aspects be used together with the aggregated data as features for downstream solutions.
[0066] FIG. 4 illustrates a set of components of an implementation of an anomaly detection module 140 (e.g., an unsupervised anomaly detection model) according to some aspects of the disclosure. The unsupervised anomaly detection model may incorporate (or be based on) a set of features 410 (e.g., corresponding to features 238). Based on the features 410 (e.g., derived by feature engineering module 130 or 230), the set of anomaly detection models 420 (e.g., including a set of anomaly detection models 420-1, 420-2, and 420-K) may be used in some aspects to generate a set of anomaly scores 430 (e.g., including a set of anomaly scores 430-1, 430-2, and 430-K) for each time point. In some aspects, the anomaly scores (e.g., anomaly scores 430-1, 430-2, or 430-K) may indicate how likely an anomaly is at each time point. For example, the anomaly scores may be defined to range from 0 to 1, with larger values representing a greater likelihood of representing an anomaly. The set of anomaly scores 430 may be used as labels and features to build a fault prediction model (e.g., a model associated with short-term fault prediction module 150 or long-term fault prediction module 170 of FIG. 1 ) in some aspects. In some aspects, multiple anomaly detection methods (represented by anomaly detection models in the set of anomaly detection models 420) may be applied to the features 410 by each anomaly detection model in the set of anomaly detection models 420 to generate anomaly scores in the set of anomaly scores 430 at each time point. The sets of anomaly scores 430 from the multiple models may be lumped together (e.g., into an ensemble anomaly score 440) in some aspects to remove bias that may be incorporated in the anomaly detection models in the set of anomaly detection models 420. A supervised surrogate model 450 may be built based on the features 410 and the ensemble anomaly scores 440 to select features 475, explain the anomaly scores (e.g., using an explainable AI model 460), and evaluate the anomaly detection model.
[0067] 1, 2, and 4 for performing anomaly detection to generate a set of anomaly scores 430 (or collectively an ensemble anomaly score 440), root causes of the anomaly scores 480, and selected features 475 are provided below. For example, the feature engineering module 130 or 230 may be used to generate a set of features 410 based on a set of physical sensor data 110 or 222 and / or virtual sensor data 226 generated by the physics-based model 120 or 220. The workflow may in some aspects include selecting multiple anomaly detection model algorithms (e.g., anomaly detection models in the set of anomaly detection models 420) and applying each selected model algorithm to the features 410 (or a subset of the features 410) to generate anomaly scores (e.g., anomaly scores 430-1, 430-2, and 430-K) from each model. For each time point, the anomaly scores 430 generated by the set of anomaly detection models 420 may be lumped together as one anomaly score.
[0068] In some aspects, the workflow may include using the features 410 as features and the ensemble anomaly scores 440 as labels to build a supervised surrogate model 450. With the supervised surrogate model 450 and the features 410, an explainable AI model 460 may be used to explain each ensemble anomaly score 440 and derive a set of root causes 480 of the ensemble anomaly scores. Each root cause (or alternatively referred to as a contributing factor) may be identified in some aspects by a feature or factor name and its weight that contributes to the ensemble anomaly score. In some aspects, an open source library may be used to explain the prediction results of a machine learning model. For example, "ELI5" (https: / / eli5.readthedocs.io / ) and "SHAP" (https: / / shap.readthedocs.io / ) are two open source libraries that may be used to explain the prediction results of a machine learning model. Such libraries are designed to explain each result every time.
[0069] The supervised surrogate model may be used in some aspects to select important features (e.g., via model-based feature selection module 470). In some aspects, model-based feature selection module 470 may use one or more feature selection techniques, such as forward selection, backward selection, or model-based feature selection (based on feature importance techniques). For example, the set of features may be associated with a calculated importance score indicating the magnitude of their contribution to at least one model utilized in the workflow (e.g., anomaly detection module 140 may be associated with the data and models shown in FIG. 4 . Thus, the output of anomaly detection module 140 may in some aspects include one or more of the individual anomaly scores 430 (or ensemble anomaly scores 440), root causes of the anomaly scores (as identified by the explainable AI model 460), and / or selected features 475 identified by model-based feature selection module 470).
[0070] In some aspects, the sensor data is collected to predict at least short-term faults. A short-term fault may be a fault that occurs in a short time period, such as a few hours or days, by design. Some examples of short-term faults may include system operation faults and / or small mechanical or electrical faults. The anomaly detection module 140 may generate an ensemble anomaly score, a set of root causes (e.g., contributing factors), as well as selected key features based on a set of historical sensor data, in some aspects. FIG. 5 is a diagram 500 illustrating a set of operations associated with short-term fault prediction. The set of operations may include a first subset of operations for model building / training (e.g., operations 502, 504, 506, and 508) and a second subset of operations for model-based estimation / prediction using a trained model on data collected in real-time (or near real-time) (e.g., operations 510, 512, and 514). In some aspects, the system may generate 502 one or more of the individual anomaly scores (or ensemble anomaly scores), root causes (contributing factors) of the anomaly scores (as identified by the explainable AI model), and / or selected features identified by the model-based feature selection module 470.
[0071] At 504, the system may define a set of parameters associated with one or more of a lookback feature window and / or a lead time window to derive features for each time point based on the anomaly scores, root causes, and features. The set of parameters associated with one or more of the lookback feature window and / or the lead time window may be defined based on domain knowledge or may be optimized based on some optimization algorithms such as grid search and random search. The set of parameters for the lookback feature window and the lead time window may include a duration of the lookback feature window in time from the current time point and a separation in time between the current time point and a potential failure occurrence time (i.e., the lead time feature window). Such parameters may, in some aspects, be based on the type of expected failure and a desired lead time for identifying a potential failure to give a user sufficient time to address the expected failure (e.g., prepare a replacement part, perform maintenance, or otherwise mitigate or avoid the failure).
[0072] Based on a set of parameters associated with the lookback feature window and / or the lead time window, the system may derive features for constructing a short-term failure prediction model at 506. For example, based on parameters associated with the lookback feature window, the system may derive features for constructing a short-term failure prediction model at 506 by calculating and / or identifying one or more of selected features, ensemble anomaly scores, and root causes thereof within a defined lookback feature window. The selected features, ensemble anomaly scores, and root causes thereof within a particular defined lookback feature window may be linked by chronological order (e.g., time point) in some aspects and used as features for generating, training, or validating a short-term failure prediction model. In some aspects, the derivation of features may also include an aggregation function based on a selected subset of the aggregate statistics. The selected aggregate statistics, in some embodiments, may include a minimum value over a time window, a maximum value over a time window, an average value over a time window, a standard deviation, a value associated with the 1st percentile, a value associated with the 99th percentile, a value associated with the 25th percentile, a value associated with the 50th percentile, a value associated with the 75th percentile, and a trend, and may be selected as described above in connection with FIG. 3.
[0073] The lead time window may further be used to generate an ensemble anomaly score for association with the lead time window (and lookback window) and the associated time point at 506. The generated set of ensemble anomaly scores may be used as a set of target data for a subsequent training operation to build / train a short-term failure prediction model. For example, the system may in some aspects define a look-ahead lead time window for each time point and use the ensemble anomaly score associated with the lead time window as a target (e.g., a ground truth for predicted values) for building a short-term failure prediction model. In some aspects, the anomaly scores are continuous values, and using continuous values as prediction targets mitigates problems associated with rare failures in the classification method.
[0074] At 508, the system may build (or train using machine learning operations) a short-term fault prediction model by the time series sequence prediction model. In some aspects, a deep learning recurrent neural network (RNN) model (e.g., LSTM, GRU) may be used to build and / or train the short-term fault prediction model. In some aspects, other methods such as autoregressive integrated moving average (ARIMA) or other suitable machine learning methods may be used. The construction of the short-term fault prediction model at 508 may represent the end of an initial model building / training subset of operations.
[0075] In operation of the system, the system calculates 510 one or more predicted failure scores based on data collected by the physical sensors and, in some embodiments, virtual sensor data. In some embodiments, the predicted failure scores may be converted to categorical failure risk levels at 512. The categorical failure risk levels may include a low risk level, a medium risk level, and a high risk level that may be more understandable to a user attempting to determine whether action should be taken based on the predicted failure scores. The conversion may be based on a set of thresholds for the predicted failure scores associated with the different risk levels. For example, in defining three different risk levels, a low risk level may be associated with a risk of having a predicted failure score of less than 0.2, a medium risk may be associated with a predicted failure score of 0.2 up to a failure prediction score of 0.6, and a high risk level may be associated with a failure prediction score of 0.6 or greater. In different embodiments, a different number of categories may be used and the labels may indicate recommended actions such as, for example, low, medium, and risk levels in the previous example, or may be substituted and / or labeled as "no action should be taken," "monitor operational status," and "repair / replace." Thus, at 514, a predicted failure score and / or categorical risk level may be reported to a user.
[0076] In some aspects, sensor data is collected to at least predict long term failures. Long term failures may be failures that occur over a long period of time, such as weeks, months, or even years. Some examples of long term failures include major mechanical failures and major electrical failures. Checking the system to identify potential long term failures can avoid major losses in terms of both assets and human safety. Figure 6 is a diagram 600 illustrating a set of operations associated with generating a long term failure prediction score to identify and predict potential long term failures according to some aspects of the present disclosure.
[0077] At 602, a separate set of anomaly detection operations may be performed for each asset in the set of assets associated with long-term failure predictions of systems or subsystems within the larger system. The set of operations may include the anomaly detection operations performed at 602. In some aspects, a first set of anomaly detection operations 602A may be performed based on sensor data, as described in connection with operation 502 of FIG. 5. Output of the anomaly detection operations performed at 602 includes an ensemble anomaly score, root causes of the ensemble anomaly scores, and selected features.
[0078] As described above in connection with FIG. 5, for each asset in a set of assets (tens to thousands of assets), one or more of the selected features within a defined time window, an ensemble anomaly score, and its root cause may be derived. Aggregate statistics similar to those described above in connection with FIG. 5 may be generated at each of the multiple timescales. The aggregate statistics at the multiple timescales may be concatenated for each asset in some aspects. The anomaly detection operations performed at 602 may further include obtaining physical design data for each asset at 602B, which in some aspects may include, but is not limited to, predicted asset life, asset material, asset type, asset model, or other relevant physical attributes.
[0079] At 604, the system may perform feature engineering operations including aggregation operations across the data in the lookback feature window for each of the multiple assets to derive a set of features and root causes (contributing factors). For example, for each asset and each set of defined time windows, the system may calculate at 604 one or more of the mean anomaly score, the variance of the anomaly scores, or other aggregate statistics, as described above. In some aspects, the aggregation reduces the "dimensionality" of the problem being modeled for long-term failure prediction, and may generate a failure prediction score that reflects or accounts for redundancy within a system or subsystem by taking into account the components of the system or subsystem.
[0080] At 606, an anomaly detection operation for the ensemble of models may be determined based on the output of the feature engineering. The feature engineering at 604 may generate a second set of system-level (aggregated) features, a set of system-level ensemble anomaly scores, and a set of system-level root causes (contributors to the anomaly scores generated by the detection models). Based on the second set of system-level (aggregated) features, the set of system-level ensemble anomaly scores, and the set of system-level root causes, the anomaly detection model at the system level (or the ensemble of the anomaly detection models) may be applied at 606 to generate a long-term failure prediction score 608. The long-term failure prediction score 608 may include a set of scores associated with different time scales. As described in connection with FIG. 5, the long-term prediction score may be categorized to allow for a more intuitive interpretation of the calculated long-term failure prediction score.
[0081] Once a short-term fault is predicted by the model, the system may be able to derive the root cause of the fault to aid in the diagnosis and repair of the short-term fault. FIG. 7 is a diagram 700 illustrating a set of operations for identifying a set of root causes associated with a short-term fault (root cause analysis of predicted short-term fault 701) according to some aspects of the present disclosure. FIG. 8 is a diagram 800 illustrating a set of operations for identifying a set of root causes associated with a long-term fault (e.g., root cause analysis of predicted long-term fault 801) according to some aspects of the present disclosure. The operations and elements of FIG. 7 and FIG. 8 overlap significantly, so the common elements are described together below. FIG. 700 illustrates that for each predicted fault represented by a predicted fault score 702 associated with a fault prediction model 704 (generated as described in connection with FIG. 5), the explainable AI model-fault prediction module 706 may identify a first feature set ("first feature set" 708) based on the fault prediction model 704 and the predicted fault score 702. In some aspects, the first feature set 708 may be a set of features identified as important features that contribute most to the predicted fault score, such features including root causes of anomaly scores 480 and selected features 475. Important features may be identified based on relative importance (e.g., a predefined number of features having the greatest impact on or contribution to the predicted fault score) or absolute importance (e.g., based on a threshold value associated with a measure of contribution to the predicted fault score 702). Thus, in some aspects, each feature is associated with a feature importance score to indicate the extent to which it contributes to the predicted fault score.
[0082] The explainable AI model-fault prediction module 706 may identify the detected anomaly scores 710 (e.g., anomaly scores 430 or ensemble anomaly scores 440) as important features of the respective predicted fault scores 702. The detected anomaly scores 710 and the supervised surrogate model-detection 712 (e.g., corresponding to the supervised surrogate model 450) may be provided to the explainable AI model-anomaly detection module 714 to derive a second set of important features (i.e., second feature set 716). For the features in the first feature set 708, each feature may be associated with a feature importance score to indicate the extent to which it contributes to the detected fault score.
[0083] To identify a set of root causes associated with the long-term faults 834, the inputs to the explainable AI model-fault prediction module 808 are slightly different from the inputs to the explainable AI model-fault prediction module 706. The operations to identify a set of root causes associated with the long-term faults 834 may include building a supervised surrogate model-prediction 806 for the long-term fault prediction model 804 (the supervised surrogate function therein is described above in connection with FIG. 4 ) by characterizing the long-term fault prediction model 804 and using anomaly scores from the long-term fault prediction model 804 as targets.
[0084] For each predicted fault, the explainable AI model-fault prediction module 808 may process the supervised surrogate model-prediction 806 and the predicted long-term fault score 802 to identify important features that contribute most to the predicted long-term fault score 802. Each feature in the first feature set 810 may, in some aspects, be associated with a feature importance score to indicate the extent to which it contributes to the predicted long-term fault score 802. The operations may include identifying a detected anomaly score 812 and a supervised surrogate model-detection 814 associated with the identified first feature set 810, and using the explainable AI model-anomaly detection module 816 to identify a second set of important features (second feature set 818). Each feature in the first feature set 810 and the second feature set 818 may, in some aspects, be associated with a feature importance score to indicate the extent to which it contributes to the detected fault score.
[0085] The operations that follow overlap significantly in short-term and long-term root cause analysis, with the understanding that the features and sensor characteristics identified using the following operations may differ due to different underlying systems and / or components being analyzed (e.g., components associated with small versus large faults).
[0086] The feature aggregation and ranking module 718 (820), in some aspects, merges features from the first feature set 708 (810) and the second feature set 716 (818) into a single set and sorts the features based on the feature importance score. In some aspects, the merging removes redundant features such that a feature that appears in both the first feature set 708 (810) and the second feature set 716 (818) may only be represented once in the merged list. For example, the feature importance score in the second feature set 716 (818) may be calculated by multiplying the feature importance score from the explainable AI model—fault prediction module 706 (808) and the feature importance score from the explainable AI model—anomaly detection module 714 (816). In some aspects, if there are overlapping features in the first Feature set 708 (810) and the second Feature set 716 (818), the system may merge the overlapping features by using an aggregated feature importance score with aggregate statistics, which may include, but are not limited to, adding the feature importance scores, taking the maximum of the feature importance scores, and taking the average of the feature importance scores. The feature aggregation and ranking module 718 (820) may sort the features in the merged list in descending order based on the feature importance scores.
[0087] For each feature in the above result set, the system may map the feature to one or more physical or virtual sensors via a module for mapping features to sensors 720 (822). For features that map directly to physical sensors, a first set of physical sensors (first physical sensor set 722 (824)) may be identified. For features that map to a set of virtual sensors 724 (826), further mapping may be performed by a module for mapping virtual sensors to physical sensors 726 (828) to identify a second set of physical sensors (second physical sensor set 728 (830)) associated with the feature. A sensor aggregation and ranking module 730 (832) may then perform merging and ranking operations on the first physical sensor set 722 (824) and the second physical sensor set 728 (830), with each physical sensor associated with an importance score based on the corresponding feature. Similar to feature aggregation and ranking, aggregating the set of identified physical sensors may involve merging the importance scores of physical sensors corresponding to multiple features based on, for example, one or more of an addition of the sensor importance scores, a maximum of the sensor importance scores, and an average of the sensor importance scores.
[0088] The physical sensors in the aggregated list may also be ranked by a sensor aggregation and ranking module 730 (832). The ranked list may include the set of physical sensors and corresponding weights indicating the magnitude of their contribution to the ensemble anomaly score sorted in descending order. The list may then be provided as a set of root causes for the predicted failure 732 (834). In some aspects, the set of first feature set 708 (810), second feature set 716 (818), and virtual sensors 724 (826) may also be output to provide further operational insight to the user.
[0089] While the above description relates to sensor data of non-moving components of the system, the method may in some aspects be extended to use such data to improve fault prediction when data on the operator is available. The operator may in some aspects include a human, a robot, or a motion profile. In some aspects, the role of the operator is important in the operation of the machine or system, and their performance directly impacts the performance of the system, which may be measured as production yield, failure rate, or user experience, and thus may provide better prediction accuracy when the role of the operator is taken into account. Specifically, in the case of failure rate, when data on the operator is collected, these may be used as additional features to build anomaly detection models, failure prediction models, and long-term failure prediction models. The data on the operator may in some aspects include, but is not limited to, one or more of operation trajectories (e.g., position, speed, and acceleration), years of experience of the operator, performance metrics of the operator, or demographic attributes of the operator.
[0090] FIG. 9 is a flow diagram illustrating a method 900 for a system for performing fault prediction operations of an industrial process. The method may be performed by a set of processing units associated with one or more computing devices associated with an industrial system including a set of components and a set of sensors for monitoring the components of the system. At 910, the system may collect (or acquire) a set of physical sensor data. The set of physical sensor data may, in some aspects, include a set of radio frequency data sampled by a sampling operation. The set of physical sensor data may include sensor data from one or more of a temperature sensor, a pressure sensor, a vibration sensor, an acoustic sensor, a motion sensor, an optical sensor, a LIDAR sensor, an IR sensor, an acceleration sensor, a gas sensor, a smoke sensor, a humidity sensor, a level sensor, an image sensor (camera), a proximity sensor, a water quality sensor, and / or a chemical sensor. For example, referring to FIG. 1, sensor data 110 may be acquired from a set of physical sensors monitoring a condition or characteristic of a component of the system.
[0091] At 920, the system may generate a virtual sensor data set by applying the physics-based model to a subset of the physical sensor data set. In some aspects, a virtual sensor is a type of software-derived information from available information that represents or is associated with data that a corresponding physical device would collect. For example, with reference to FIG. 1 and FIG. 2A, the physics-based model 120 or 220 (including the physics-based model 224) may take the physical sensor data 110 or 222 and generate a set of virtual sensor data 226 based on the physics-based model (e.g., the physics-based model 224). The physics-based model 224 may in some aspects be a representation of laws of nature that inherently incorporate concepts of time, space, causality, and generalizability. These laws of nature, in some aspects, define how physical, chemical, biological, and geological processes evolve. The physics-based model 224 may in some aspects be a function that takes multiple inputs (e.g., the physical sensor data 222) and generates multiple outputs (e.g., the virtual sensor data 226). The inputs may come from predefined profiles (such as motion profiles) at design time or from physical sensors (e.g., physical sensor data 222) at operational time. The outputs (e.g., virtual sensor data 226) may in some aspects include multiple variables and may represent a set of virtual sensors.
[0092] At 930, the system may identify a first set of features from the physical sensor data set and the virtual sensor data set by performing at least one of a sampling operation, an aggregation operation, or a feature derivation operation on the physical sensor data set and the virtual sensor data set based on the optimized sampling rate or the optimized aggregation statistics. The physical sensor data set may include a high frequency data set in some aspects, and the optimized sampling rate may be associated with a sampling operation to reduce the amount of data. For example, with reference to FIG. 2B and FIG. 4, the feature engineering module 230 may generate the set of features 238 or 410 using one or more of the sampling and aggregation module 234 and / or the feature derivation module 236 applied to the sensor data 232. As described in connection with FIG. 3, the optimized sampling rate and / or the optimized aggregation statistics may be based on one or more optimization algorithms. For example, at least one of the optimized sampling rate or the optimized aggregation statistics may be calculated using one or more of a Bayesian optimization, a grid search, or a random search in some aspects. As described above, in some aspects, the first optimized sampling rate or optimized aggregate statistic is calculated based on domain knowledge, and the second optimized sampling rate or optimized aggregate statistic is calculated based on one or more of Bayesian optimization, grid search, or random search. In some aspects, optimizing the first optimized sampling rate or optimized aggregate statistic and the second optimized sampling rate or optimized aggregate statistic includes cycling between or iterating between these optimizations based on results of previous optimizations, as described in connection with FIG. 3. In some aspects, the optimized sampling and aggregation operations may be implemented by a first set of machine training models, e.g., trained by the method described above in connection with FIG. 3.
[0093] At 940, the system may identify a set of anomaly detection scores, a first set of contributors to the set of anomaly detection scores, and a second set of features having feature importance scores that are equal to or greater than a threshold based on a machine training model (e.g., a trained sampling and aggregation model, which may correspond to an anomaly detection model) applied to the first set of features. The first set of machine training models may in some aspects include a set of anomaly detection modules that generate the set of anomaly detection scores. The set of anomaly detection models may in some aspects be trained based on the set of physical sensor data collected at 910 and the set of virtual sensor data generated at 920. For example, referring to FIG. 4 , based on the set of anomaly detection models 420 applied to features 410, the ensemble anomaly scores 440, the supervised surrogate model 450, the explainable AI model 460, and the model-based feature selection 470, the system may identify (1) a set of anomaly detection scores 430 (e.g., anomaly scores 430-1, 430-2, and 430-K or collectively the ensemble anomaly scores 440), (2) root causes of the anomaly scores 480 as a first set of anomaly score contributors 480 (e.g., a first set of contributors to the set of anomaly detection scores), and (3) selected features 475 (e.g., a second set of features having feature importance scores above a threshold).
[0094] At 950, the system may generate at least one failure prediction model based on the first set of features and the first set of machine training models. The at least one failure prediction model, in some aspects, may include at least one short-term failure prediction model based on (1) the set of anomaly detection scores, (2) sampling and / or aggregation operations, or (3) machine learning operations applied to the first set of identified features. The machine learning operations used to build and / or train the short-term failure predictions, in some aspects, may include one or more deep learning RNN models (e.g., LSTM, GRU), ARIMA, or other suitable machine learning models. The at least one short-term failure prediction model, in some aspects, may include a short-term failure prediction model generated for each individual asset in the set of assets associated with the system. For example, referring to FIG. 5 , the system may generate (e.g., build / train) 508 a short-term fault prediction model based on (1) anomaly detection model set-based operation 502, (2) parameter optimization operation set-based operation 504, and (3) feature derivation and transformation operation set-based operation 506.
[0095] In some aspects, the at least one failure prediction model may include at least one long-term failure prediction model. Generating the at least one long-term failure prediction model may, in some aspects, be further based on a second set of machine training models. The second set of machine training models may, in some aspects, be based on the first set of machine training models. For example, with reference to FIG. 6, the system may generate the at least one long-term failure prediction model based on a further set of anomaly detection models (e.g., a second set of machine training models) applied at 606 to a set of data generated based on the first set of machine training models applied to each of the multiple assets at 602.
[0096] At 960, the system may apply at least one failure prediction model to the set of physical sensor data and the set of virtual sensor data to calculate a likelihood score for a predicted failure of the asset. Further operations that may be performed are indicated by the letter "A".
[0097] FIG. 10 is a flow diagram illustrating a method 1000 of identifying a set of root causes (e.g., contributing factors) for failure prediction according to some aspects of the disclosure. As indicated by the letter "A", the method of FIG. 10 may be performed after the operations described in connection with the method illustrated in FIG. 9. The method 1000 is performed in some aspects by a root cause analysis module (e.g., root cause analysis for short-term faults module 160 or root cause analysis for long-term faults module 180 of FIG. 1). At 1010, the root cause analysis module may in some aspects identify a set of contributing features of at least one failure prediction model by applying a first explanatory model for at least one failure prediction model to the set of predicted fault scores and applying a second explanatory model for the first set of machine training models. In some aspects, the term contributing features may refer to abstract features identified by a model in a system, while the term contributing factors may relate to physical components of the system identified as contributing to the failure prediction model. For example, with reference to Figures 7 and 8, the explainable AI model-fault prediction module 706 (or 808) may obtain a set of predicted short-term (or long-term) fault scores 702 (or 802) to identify a first feature set 708 (810) associated with a detected anomaly score 710 (812). The explainable AI model-anomaly detection module 714 (816) may also be used to identify a second feature set 716 (818). The first feature set 708 (810) and the second feature set 716 (818) may together constitute a set of contributors of at least one fault prediction model.
[0098] At 1020, the root cause analysis module may, in some embodiments, calculate a second importance score for each contributing feature in the set of contributing features. The calculated second importance score, in some embodiments, reflects a contribution to one or more fault prediction scores or anomaly scores. For example, with reference to FIG. 7 and FIG. 8, the explainable AI model-fault prediction module 706 (or 808) and the explainable AI model-anomaly detection module 714 (816) may calculate or output a weight associated with each contributing feature. The root cause analysis module (e.g., root cause analysis of predicted short-term failures 701 or root cause analysis of predicted long-term failures 801) may generate a list of features associated with corresponding weights, and the feature aggregation and ranking module 718 (or 820) may generate a ranked list of features and aggregated weights (e.g., features contributing to multiple anomaly scores and / or predictions).
[0099] At 1030, the root cause analysis module may, in some aspects, map each contributing feature in the set of contributing features to one or more contributing physical sensors. The mapping may include a mapping of features to physical sensors and a mapping of features to virtual sensors. In addition, the mapping of each contributing feature in the set of contributing features to one or more contributing physical sensors may also include a mapping of virtual sensors to one or more physical sensors. For example, with reference to FIGS. 7 and 8, the module mapping features to sensor 720 (822) may map contributing factors in a first set of factors to a set of physical sensors 722 (824) and 728 (830), e.g., via an intermediate mapping to a set of virtual sensors 724 (826), which are then mapped to a set of physical sensors 728 (830).
[0100] At 1040, the root cause analysis module may, in some aspects, calculate a third importance score for each of the one or more contributing physical sensors based on the second importance score of at least one contributing feature in the set of contributing features mapped to the physical sensor of the one or more contributing physical sensors. The calculation of the third importance score may, in some aspects, be based on aggregating weights associated with the same physical sensor based on different features or mappings. For example, with reference to FIGS. 7 and 8, the sensor aggregation and ranking module 730 (832) may perform a merge operation on the first physical sensor set 722 (824) and the second physical sensor set 728 (830), each physical sensor associated with an importance score based on one or more corresponding features. Similar to feature aggregation and ranking, aggregating the set of identified physical sensors may merge the importance scores of physical sensors corresponding to multiple features based on one or more of, for example, an addition of the sensor importance scores, a maximum of the sensor importance scores, and an average of the sensor importance scores.
[0101] Finally, at 1050, the root cause analysis module may, in some aspects, identify a subset of the contributing physical sensors of the one or more contributing physical sensors as a second set of contributing factors of at least one fault prediction model based on the third importance score. For example, with reference to FIGS. 7 and 8, the sensor aggregation and ranking module 730 (832) may perform a ranking operation on the merged list of contributing factors based on the first physical sensor set 722 (824) and the second physical sensor set 728 (830). The physical sensors in the aggregated list may be ranked by the sensor aggregation and ranking module 730 (832). The ranked list may include a set of physical sensors and corresponding weights indicating the magnitude of their contribution to the ensemble anomaly score sorted in descending order.
[0102] As presented in this disclosure, the system may provide one or more of the following benefits. For example, the disclosure described above introduces an automated, unsupervised, data-driven solution for fault / failure control in industrial systems for one or more of short-term faults (in hours or days) or long-term faults (in weeks, months, or years). The disclosure described above provides a solution that can predict faults and identify the root cause of the fault with some lead time so that an operator / technician may have enough time to respond and may have further root cause information to facilitate diagnosis of the fault. The disclosure described above reduces or eliminates the need for (manual) labeling data. While in some aspects the disclosure relates to performing the above-mentioned operations on historical sensor data, historical fault data may be optional in some aspects. The disclosure also introduces a semi-empirical approach through a combination of physics-based and machine learning models, and optimization strategies for sampling rate and aggregation statistics for high frequency sensor data.
[0103] In some aspects, the proposed anomaly prediction method performs well in predicting rare failure events a priori with advanced deep learning sequence prediction power and continuous values as targets. The present disclosure also relates to root causes at several levels (detected anomaly score level, feature level, raw sensor data level) identified for predicted failures through a chain of explainable AI techniques and aggregation / ranking algorithms.
[0104] 11 illustrates an exemplary computing environment having an exemplary computing device suitable for use in some exemplary implementations. The computing device 1105 in the computing environment 1100 may include one or more processing units, cores, or processors 1110, memory 1115 (e.g., RAM, ROM, and / or the like), internal storage 1120 (e.g., magnetic, optical, solid-state storage, and / or organic), and / or IO interface 1125, any of which may be connected over a communication mechanism or bus 1130 to communicate information or may be incorporated within the computing device 1105. The IO interface 1125 may also be configured to receive images from a camera or provide images to a projector or display, depending on the desired implementation.
[0105] The computing device 1105 may be communicatively connected to an input / user interface 1135 and an output device / interface 1140. One or both of the input / user interface 1135 and the output device / interface 1140 may be wired or wireless interfaces and may be removable. The input / user interface 1135 may include any physical or virtual device, component, sensor, or interface that may be used to provide input (e.g., buttons, touch screen interfaces, keyboards, pointing / cursor controls, microphones, cameras, Braille, motion sensors, accelerometers, optical readers, and / or the like). The output devices / interface 1140 may include displays, televisions, monitors, printers, speakers, Braille, or the like. In some example implementations, the input / user interface 1135 and the output device / interface 1140 may be incorporated with or physically connected to the computing device 1105. In other example implementations, other computing devices may function as or provide the functionality of input / user interface 1135 and output device / interface 1140 for computing device 1105 .
[0106] Examples of computing devices 1105 may include, but are not limited to, highly mobile devices (e.g., smart phones, devices in vehicles or other machines, devices carried by humans and animals, and the like), mobile devices (e.g., tablets, notebooks, laptops, personal computers, portable televisions, radios, and the like), and devices not designed for mobility (e.g., desktop computers, other computers, information kiosks, televisions, radios, and the like having one or more processors embedded therein and / or connected thereto).
[0107] Computing device 1105 may be communicatively connected (e.g., via IO interface 1125) to external storage 1145 and a network 1150 for communicating with any number of networked components, devices, and systems, including one or more computing devices of the same or different configurations. Computing device 1105 or any connected computing device may function as, provide functionality for, or be referred to as a server, client, thin server, general purpose machine, special purpose machine, or another label.
[0108] IO interface 1125 may include, but is not limited to, wired and / or wireless interfaces using any communication or IO protocol or standard (e.g., Ethernet, 902.11x, Universal Serial Bus, WiMax, modem, cellular network protocols, and the like) to communicate information to and / or from at least all connected components, devices, and networks in computing environment 1100. Network 1150 may be any network or combination of networks (e.g., the Internet, a local area network, a wide area network, a telephone network, a cellular network, a satellite network, and the like).
[0109] The computing device 1105 may use and / or communicate using computer usable or computer readable media, including transitory media and non-transitory media. Transitory media include transmission media (e.g., metallic cables, optical fibers), signals, carrier waves, and the like. Non-transitory media include magnetic media (e.g., disks and tapes), optical media (e.g., CD-ROMs, digital video disks, Blu-ray disks), solid-state media (e.g., RAM, ROM, flash memory, solid-state storage), and other non-volatile storage or memory.
[0110] The computing device 1105 can be used to implement techniques, methods, applications, processes, or computer-executable instructions in some exemplary computing environments. The computer-executable instructions can be obtained from a transitory medium and stored on and obtained from a non-transitory medium. The executable instructions can be from one or more of any programming, scripting, and machine language (e.g., C, C++, C#, Java, Visual Basic, Python, Perl, JavaScript, and others).
[0111] The processor 1110 may execute under any operating system (OS) (not shown) in a native or virtual environment. Along with the OS and other applications (not shown), one or more applications may be deployed, including a logic unit 1160, an application programming interface (API) unit 1165, an input unit 1170, an output unit 1175, and an inter-unit communication mechanism 1196 for different units to communicate with each other. The described units and elements may be modified in design, function, configuration, or implementation, and are not limited to the description provided. The processor 1110 may be in the form of a hardware processor, such as a central processing unit (CPU), or a combination of hardware and software units.
[0112] In some example implementations, information or instructions to execute, once received by the API unit 1165, may be communicated to one or more other units (e.g., logic unit 1160, input unit 1170, output unit 1175). In some cases, logic unit 1160 may be configured to control information flow between units and control services provided by API unit 116, input unit 1170, output unit 1175 in some example implementations described above. For example, the flow of one or more processes or implementations may be controlled by logic unit 1160 alone or in cooperation with API unit 1165. Input unit 1170 may be configured to obtain inputs for computations described in the example implementations, and output unit 1175 may be configured to provide outputs based on computations described in the example implementations.
[0113] The processor 1110 may be configured to collect a physical sensor data set. The processor 1110 may also be configured to generate a virtual sensor data set by applying a physics-based model to a subset of the physical sensor data set. The processor 1110 may be further configured to identify a first set of features from the physical sensor data set and the virtual sensor data set by performing at least one of a sampling operation, an aggregation operation, or a feature derivation operation on the physical sensor data set and the virtual sensor data set based on an optimized sampling rate or an optimized aggregation statistic. The processor 1110 may be further configured to identify a set of anomaly detection scores, a first set of contributors to the set of anomaly detection scores, and a first set of features having a feature importance score that is equal to or greater than a threshold using a first set of machine training models applied to the first set of features. The processor 1110 may be further configured to generate at least one fault prediction model based on the first set of machine training models. The processor 1110 may also be configured to apply at least one fault prediction model to the set of physical sensor data and the set of virtual sensor data to calculate a likelihood score for a predicted fault of the asset. The processor 1110 may also be configured to identify a set of contributing factors for the at least one fault prediction model by applying a first explanatory model for the at least one fault prediction model to the set of predicted fault scores and applying a second explanatory model for the first set of machine training models. The processor 1110 may also be configured to calculate a second importance score for each contributing feature in the set of contributing features. The processor 1110 may also be configured to map each contributing feature in the set of contributing features to one or more contributing physical sensors. The processor 1110 may also be configured to calculate a third importance score for each of the one or more contributing physical sensors based on the second importance score of at least one contributing feature in the set of contributing features mapped to that physical sensor of the one or more contributing physical sensors.The processor 1110 may also be configured to identify a subset of the contributing physical sensors of the one or more contributing physical sensors as a second set of contributing factors of the at least one fault prediction model based on the third importance score.
[0114] Some portions of the detailed description are presented in terms of algorithms and symbolic representations of operations within a computer. These algorithmic descriptions and symbolic representations are the means used by those skilled in the data processing arts to convey the substance of their innovations to others skilled in the art. An algorithm is a sequence of defined steps leading to a desired end state or result. In one implementation, the steps performed require physical manipulations of tangible quantities to achieve a tangible result.
[0115] Unless otherwise specified, as is apparent from the description, the description utilizing terms such as "processing," "computing," "calculating," "determining," "displaying," and the like throughout the description will be understood to include the actions and processes of a computer system or other information processing device that manipulates and converts data represented as physical (electronic) quantities in the registers and memory of the computer system to other data similarly represented as physical quantities in the memory or registers of the computer system, or other information storage, transmission, or display device.
[0116] The implementations may also relate to an apparatus for performing the operations of the present specification. The apparatus may be specially constructed for the required purposes, or may include one or more general-purpose computers selectively activated or reconfigured by one or more computer programs. Such computer programs may be stored in a computer-readable medium, such as a computer-readable storage medium or a computer-readable signal medium. The computer-readable storage medium may include tangible media, such as, but not limited to, optical disks, magnetic disks, read-only memories, random access memories, solid-state devices and drives, or any other type of tangible or non-transitory medium suitable for storing electronic information. The computer-readable signal medium may include media such as carrier waves. The algorithms and displays presented herein are not inherently related to any particular computer or other apparatus. The computer programs may include pure software implementations that include instructions to perform the operations of the desired implementation.
[0117] Various general-purpose systems may be used with the programs and modules according to the examples herein, or it may prove convenient to construct more specialized apparatus to perform the desired method steps. Additionally, the implementations are not described with reference to a particular programming language. It will be appreciated that a variety of programming languages may be used to implement the teachings of the implementations described herein. Instructions in a programming language may be executed by one or more processing units, such as a central processing unit (CPU), processor, or controller.
[0118] As is known in the art, the operations described above may be performed by hardware, software, or some combination of software and hardware. Various aspects of the implementations may be implemented using circuits and logic (hardware), while other aspects may be implemented using instructions stored on a machine-readable medium (software) that, when executed by a processor, cause the processor to perform a method for performing the implementation of the present application. Furthermore, some implementations of the present application may be performed exclusively in hardware, while other implementations may be performed exclusively in software. Furthermore, the various functions described may be performed in a single unit or distributed among several components in any number of ways. When performed by software, the method may be executed by a processor, such as a general purpose computer, based on instructions stored on a computer-readable medium. If desired, the instructions may be stored on the medium in compressed and / or encrypted format.
[0119] Additionally, other implementations of the present application will be apparent to those skilled in the art upon consideration of the specification and practice of the techniques of the present application. Various aspects and / or components of the described implementations may be used alone or in any combination. It is intended that the specification and implementations be considered as examples only, with a true scope and spirit of the present application being indicated by the following claims.
Claims
1. collecting a set of physical sensor data; generating a virtual sensor data set by applying a physics-based model to a subset of the physical sensor data set; identifying a first set of features from the physical sensor data sets and the virtual sensor data sets by performing at least one of a sampling operation, an aggregation operation, or a feature derivation operation on the physical sensor data sets and the virtual sensor data sets based on an optimized sampling rate or optimized aggregation statistics; identifying a set of anomaly-detection scores, a first set of contributors to the set of anomaly-detection scores, and a second set of features having feature importance scores that are equal to or greater than a threshold using a first set of machine-trained models applied to the first set of features; generating at least one fault prediction model based on the first set of machine trained models; applying the at least one failure prediction model to the set of physical sensor data and the set of virtual sensor data to calculate a likelihood score for a predicted failure of an asset; The method includes:
2. 2. The method of claim 1, further comprising identifying a second set of contributing factors for the at least one failure prediction model by applying a first explanatory model for the at least one failure prediction model to a set of predicted failure scores generated by the at least one failure prediction model.
3. The second set of contributing factors includes a set of contributing physical sensors, and identifying the second set of contributing factors includes: identifying a set of contributing features of the at least one fault prediction model by applying the first explanatory model for the at least one fault prediction model to the set of predicted fault scores and applying a second explanatory model for the first set of machine trained models; calculating a second importance score for each contributing feature in the set of contributing features; mapping each contributing feature in the set of contributing features to one or more sensors in the set of contributing physical sensors; calculating, for each physical sensor in the set of contributing physical sensors, a third importance score based on the second importance score of at least one contributing feature in the set of contributing features mapped to the physical sensor in the set of contributing physical sensors; The method of claim 2 , comprising:
4. The method of claim 1 , wherein the physical sensor data set comprises a high frequency data set sampled by the sampling operation.
5. The method of claim 1 , wherein the first set of machine-trained models comprises a set of anomaly detection models.
6. The method of claim 5 , wherein the set of anomaly detection models is trained based on the set of physical sensor data and the set of virtual sensor data.
7. 2. The method of claim 1 , wherein the at least one failure prediction model comprises at least one short-term failure prediction model based on machine learning operations applied to a set of anomaly-detection scores, the first set of contributing factors to the set of anomaly-detection scores, and a second set of features generated by the first set of machine training models.
8. The method of claim 7 , wherein the at least one short-term failure prediction model comprises a short-term failure prediction model generated for each individual asset.
9. 2. The method of claim 1, wherein the at least one failure prediction model includes at least one long-term failure prediction model, and generating the at least one long-term failure prediction model is further based on a second set of machine training models, the second set of machine training models being based on the first set of machine training models.
10. The method of claim 1 , wherein at least one of the optimized sampling rate or the optimized summary statistic is calculated using one or more of a Bayesian optimization, a grid search, or a random search.
11. 2. The method of claim 1 , wherein a first one of the optimized sampling rate or the optimized aggregate statistic is calculated based on domain knowledge, and a second one of the optimized sampling rate or the optimized aggregate statistic is calculated based on one or more of a Bayesian optimization, a grid search, or a random search.
12. Memory, When the program stored in the memory is executed, collecting a set of physical sensor data; generating a virtual sensor data set by applying a physics-based model to a subset of the physical sensor data set; identifying a first set of features from the physical sensor data set and the virtual sensor data set by performing at least one of a sampling operation, an aggregation operation, or a feature derivation operation on the physical sensor data set and the virtual sensor data set based on an optimized sampling rate or optimized aggregation statistics; identifying a set of anomaly-detection scores, a first set of contributors to the set of anomaly-detection scores, and a second set of features having feature importance scores that are equal to or greater than a threshold using a first set of machine-trained models applied to the first set of features; generating at least one fault prediction model based on the first set of machine trained models; applying the at least one failure prediction model to the set of physical sensor data and the set of virtual sensor data to calculate a likelihood score for a predicted failure of an asset; a set of processors coupled to the memory configured to An apparatus comprising:
13. The at least one processor identifying a set of contributing features of the at least one fault prediction model by applying a first explanatory model for the at least one fault prediction model to a set of predicted fault scores and applying a second explanatory model for the first set of machine training models; calculating a second importance score for each contributing feature in the set of contributing features; mapping each contributing feature in the set of contributing features to one or more contributing physical sensors; calculating, for each of the one or more contributing physical sensors, a third importance score based on the second importance score of at least one contributing feature in the set of contributing features mapped to the physical sensor of the one or more contributing physical sensors; identifying a subset of contributing physical sensors of the one or more contributing physical sensors as a second set of contributing factors of the at least one fault prediction model based on the third importance score; The apparatus of claim 12 , further configured to:
14. The apparatus of claim 12 , wherein the physical sensor data set comprises a high frequency data set sampled by the sampling operation.
15. The apparatus of claim 12 , wherein the first set of machine-trained models comprises a set of anomaly detection models.
16. The apparatus of claim 15 , wherein the set of anomaly detection models is trained based on the set of physical sensor data and the set of virtual sensor data.
17. 13. The apparatus of claim 12, wherein the at least one failure prediction model comprises at least one short-term failure prediction model based on machine learning operations applied to a set of anomaly-detection scores, the first set of contributing factors to the set of anomaly-detection scores, and a second set of features generated by the first set of machine training models.
18. 20. The apparatus of claim 17, wherein the at least one short-term failure prediction model comprises a short-term failure prediction model generated for each individual asset.
19. 13. The apparatus of claim 12, wherein the at least one failure prediction model includes at least one long-term failure prediction model, and generating the at least one long-term failure prediction model is further based on a second set of machine training models, the second set of machine training models being based on the first set of machine training models.
20. 13. The apparatus of claim 12, wherein at least one of the optimized sampling rate or the optimized aggregate statistic is calculated using one or more of a Bayesian optimization, a grid search, or a random search, a first of the optimized sampling rate or the optimized aggregate statistic being based on domain knowledge, and a second of the optimized sampling rate or the optimized aggregate statistic being calculated based on one or more of the Bayesian optimization, the grid search, or the random search.
Citation Information
Patent Citations
Abnormality detection system, support device and model generation method
JP2019159902A
Multi-stage failure analysis and prediction
US20140281713A1
Abnormality detection system, support device, and model generation method
US20190286096A1