Real-time Detection, Prediction, and Repair of Sensor Failures through a Data-driven Approach

An automated data-driven method for IoT and OT systems addresses the challenge of timely sensor failure detection by identifying correlated sensors, detecting faults, and implementing repair strategies, thereby ensuring continuous operation and reducing costs.

JP2025517109AActive Publication Date: 2025-06-03HITACHI VANTARA LLC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2024563988
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2022-04-29
Publication Date
2025-06-03
Estimated Expiration
2042-04-29

AI Technical Summary

Technical Problem

Existing IoT and OT systems face challenges in timely detection and management of sensor failures, leading to inaccurate data and increased costs due to manual inspections and potential system downtime.

Method used

An automated data-driven approach that includes receiving sensor data from multiple related sensors, identifying correlated sensors, detecting faults, and implementing repair strategies to maintain system operation and accuracy.

Benefits of technology

This approach enables real-time fault detection and prediction, reducing manual inspection errors and costs, ensuring continuous system operation, and implementing systematic fault tolerance strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025517109000001_ABST
    Figure 2025517109000001_ABST
Patent Text Reader

Abstract

A method for real-time detection, prediction, and repair of sensor failures may include receiving sensor data from a plurality of related sensors. The method may include identifying a set of correlated sensors in the plurality of related sensors for a first sensor in the plurality of related sensors. The method may further include detecting a failure in the first sensor based on at least one of the sensor data received from the first sensor, the sensor data received from the set of correlated sensors, and the sensor data received from other sensors. The method may further include implementing a repair strategy based on the predicted failure of the sensor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to the Internet of Things (IoT) and Operational Technology (OT) areas.

Background Art

[0002] IoT and OT offer great potential to change system functions and business operation methods by efficiently monitoring and automating systems without the need for human interaction or involvement. IoT and OT systems rely on large amounts of data collected by one or more sensors in some applications to automate system operation and decision-making. A sensor may, in some embodiments, be a device that responds to an input from the physical world, captures the input, and transmits them to a storage device.

[0003] As used herein, a sensor is a device designed to respond to and / or monitor a particular type of condition in the physical world and then generate a signal (usually an electrical signal) that can represent the magnitude of the monitored condition. As the application of IoT devices and OT expands, different types of sensors are used, resulting in different types of data for analysis and processing. In some embodiments, the sensor may include any one or more of a temperature sensor, a pressure sensor, a vibration sensor, an acoustic sensor, a motion sensor, a level sensor, an image sensor, a proximity sensor, a water quality sensor, a chemical sensor, a gas sensor, a smoke sensor, an infrared (IR) sensor, an acceleration sensor, a gyro sensor, a humidity sensor, an optical sensor, and / or a Light Detection and Ranging (LIDAR) sensor.

[0004] Sensor data collected from different types of sensors may be represented differently. For example, some sensors may be analog sensors that capture continuous values and attempt to identify every nuance of what is being measured, or digital sensors that may use sampling to encode what is being measured. As a result, the captured data can be either "analog data" or "digital data". Thus, the data may be numerical, an image, or a video. In addition, some sensors collect data in a streaming manner and use time-series data to represent the values collected. Other sensors may collect data at isolated points in time.

[0005] IoT and OT industrial systems rely, in some aspects, on sensors that function to monitor the system and collect accurate data for processing, analysis, and modeling in a set of downstream applications. The quality of data from sensors plays a fundamental role in the IoT and OT domains in some aspects. Due to the nature of deployment (which can be in the field and / or harsh environments) and limitations of low-cost components, sensors can tend to fail. In some aspects, a significant portion of the failures can result from drifts and sudden failures in the sensing components of the sensors, leading to serious data inaccuracies. As a result, IoT sensors can drift, stop functioning, become unreliable, and output data that can cause misunderstandings after operating for some time. In an IoT / OT system, sensors may be installed on assets and connected through a network to storage and / or computing servers for data collection and processing. Any part of the hardware or software used to support the operation of the sensors can malfunction and result in incorrect sensor readings. Failures can occur at the root layer (the sensor), the network layer (network connectivity), the computing layer, or the storage layer. To operate the IoT / OT system accurately and continuously, it is useful to detect failures at all layers, but this disclosure focuses on failures in the sensors, including the direct link to the sensors (which is part of the network layer).

[0006] Currently, in schedule-based inspections, a failed sensor cannot be captured in a timely manner, and unnecessary inspections incur additional costs. Such manual inspections are error-prone and may take a long time. This disclosure addresses an automated data-driven approach for detecting faults in sensors and further predicting faults in sensors. Additionally, root cause analysis is performed for individual faults, and a systematic fault tolerance strategy can be designed to enable the IoT system to continue operating without interruption despite faults in one or more sensors. In some embodiments, based on the detected faulty sensors, the system can identify and execute repair actions to repair or replace the sensors in order to avoid any incorrect decisions based on readings from such faulty sensors. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM

[0007] Exemplary implementations described herein include innovative methods. The method may include receiving sensor data from a plurality of related sensors. The method may further include identifying a set of correlated sensors in the plurality of related sensors for a first sensor in the plurality of related sensors. The method may include detecting a fault in the first sensor based on at least one of the sensor data received from the sensor, the sensor data received from the set of correlated sensors, and the sensor data received from other sensors. The method may further include implementing a repair strategy based on the detected fault in the first sensor.

[0008] The exemplary implementations described herein include an innovative computer-readable medium storing computer-executable code. The computer-executable code may include instructions for receiving sensor data from a plurality of related sensors. The computer-executable code may also include instructions for identifying a set of correlated sensors in the plurality of related sensors for a first sensor in the plurality of related sensors. The computer-executable code may further include detecting a failure in the first sensor based on at least one of the sensor data received from the sensor, the sensor data received from the set of correlated sensors, and the sensor data received from other sensors. The computer-executable code may also include instructions for implementing a repair strategy based on the detected failure of the first sensor.

[0009] The exemplary implementations described herein include an innovative apparatus. The apparatus may include a memory and at least one processor configured to collect a set of physical sensor data. The at least one processor may also be configured to receive sensor data from a plurality of related sensors. The at least one processor may further be configured to identify a set of correlated sensors in the plurality of related sensors for a first sensor in the plurality of related sensors. The at least one processor may also be configured to detect a failure in the first sensor based on at least one of the sensor data received from the sensor, the sensor data received from the set of correlated sensors, and the sensor data received from other sensors. The at least one processor may be configured to implement a repair strategy based on the detected failure of the first sensor.

[0010] The exemplary implementations described herein include innovative devices. The device may include means for receiving sensor data from a plurality of related sensors. The device may further include means for identifying a set of correlated sensors in the plurality of related sensors for a first sensor among the plurality of related sensors. The device may also include means for detecting a fault in the first sensor based on at least one of the sensor data received from the sensor, the sensor data received from the set of correlated sensors, and the sensor data received from other sensors. The device may further include means for implementing a repair strategy based on the detected fault in the first sensor.

Brief Description of the Drawings

[0011]

Figure 1

[0012]

Figure 2A

[0013]

Figure 2B

[0014]

Figure 3A

[0015]

Figure 3B

[0016]

Figure 4A

[0017]

Figure 4B

[0018]

Figure 5

[0019]

Figure 6

[0020]

Figure 7

[0021]

Figure 8

[0022]

Figure 9

[0023] The following detailed description provides details of the drawings and implementation examples of the present application. Reference numerals and descriptions of redundant elements between the drawings are omitted for clarity. The terms used throughout the description are provided as examples and are not intended to be limiting. For example, the use of the term "automated" may include fully automated implementation forms or semi-automated implementation forms that include user or administrator control for certain aspects of the implementation form, depending on the desired implementation form of those skilled in the art practicing the implementation forms of the present application. The selection may be made by the user via a user interface or other input means, or may be implemented by a desired algorithm. The implementation examples described herein may be used alone or in combination, and the functions of the implementation examples may be implemented by any means according to the desired implementation form.

[0024] In the present disclosure, systems, devices, and methods are presented for addressing problems associated with conventional sensor fault detection techniques. For example, conventional approaches may not be able to detect faults in sensors in a timely manner, which can lead to incorrect readings, damage to the system, generation of inaccurate insights, and incorrect decisions. Further, conventional approaches may detect faults after they have already occurred and thus may not be able to repair or avoid the faults beforehand. Some approaches use conventional time series prediction techniques for predicting anomalies in data, but existing approaches may not be able to distinguish sensor faults from operational anomalies in the underlying system. Manual inspection of IoT sensors is error-prone and time-consuming. For example, schedule-based inspection of IoT sensors may not capture faulty sensors in a timely manner, thus posing a risk of obtaining incorrect sensor readings, and an overly aggressive inspection schedule designed to mitigate the risk of incorrect sensor readings may be associated with additional unnecessary costs.

[0025] The present disclosure presents systems, apparatuses, and methods for providing techniques related to detecting faults in sensors associated with a system (e.g., IoT sensors associated with an industrial and / or manufacturing system and / or process). The method may include receiving sensor data from a plurality of related sensors. The method may further include, for a first sensor among the plurality of related sensors, identifying a set of correlated sensors among the plurality of related sensors. The method may also include detecting a fault in the first sensor based on at least one of the sensor data received from the sensor, the sensor data received from the set of correlated sensors, and the sensor data received from other sensors. The method may further include implementing a repair strategy based on the detected fault in the first sensor.

[0026] Generally, to address some of the problems identified above, the method may involve one or more of critical sensor identification, fault tolerance identification, fault detection, fault prediction, fault repair, and / or fault tolerance. For example, critical sensor identification may include identifying one or more critical sensors based on domain knowledge, data analysis, or downstream tasks. A critical sensor may be a sensor that captures critical data for monitoring the health of the underlying system and / or critical data that is used for at least one of identifying a repair strategy, deriving business insights, or constructing solutions for problems related to a set of related downstream tasks such as anomaly detection, fault prediction, remaining useful life prediction, etc.

[0027] Fault tolerance identification may, in some embodiments, include identifying a set of one or more correlated sensors for each critical sensor (e.g., the target sensor). The identified correlated sensors may include one or more sensors that capture signals that are similar or highly correlated based on a similarity metric, and some approaches may use one or more similarity scores (e.g., similarity scores associated with a similarity metric) between the data of two sensors. Fault detection may, in some embodiments, include detecting a fault in one or more sensors based on data from physical sensors and / or data associated with virtual sensors (e.g., predicted data for virtual sensors based on relevant data from one or more physical sensors processed using one or more physics-based models). Fault detection may, in some embodiments, include one or more of univariate anomaly detection, bivariate anomaly detection, and / or multivariate anomaly detection approaches, and may further involve an ensemble algorithm based on one or more of the univariate, bivariate, or multivariate anomaly detection approaches.

[0028] In some embodiments, fault prediction may include predicting a fault in one or more sensors (e.g., critical sensors). Fault prediction may be based on a deep learning recurrent neural network (RNN) model that uses sensor data from at least the critical sensors and additional sensors (e.g., a set of relevant and / or correlated sensors). Based on the fault prediction, in some embodiments, the method may include identifying fault repair and / or fault tolerance actions. Fault repair actions (e.g., repairing or replacing a sensor predicted to fail) may be based on a root cause analysis related to the fault prediction and may further be based on domain knowledge indicating one or more repair strategies based on the results of the root cause analysis. In some embodiments, the fault prediction for a particular target sensor enables the system (or method) to identify (or obtain the identity of) a set of correlated sensors that can be used in place of the particular target sensor until the sensor is repaired or replaced.

[0029] As will be described in more detail below, the apparatus and methods described herein may provide a data-driven approach for fault detection in sensors that can distinguish between faults in sensors and operational faults in the underlying system based on a novel combination of univariate anomaly detection and bivariate / multivariate anomaly detection. In some aspects, the data-driven approach uses data from a plurality of sensors associated with the system (e.g., a set of correlated and / or related sensors) to detect and / or predict faults in a particular sensor of interest. A data-driven approach that includes one or more of critical sensor identification, fault tolerance identification, fault detection, fault prediction, fault repair, and / or fault tolerance may, in some aspects, (1) avoid damage to an unmonitored underlying system based on a faulty sensor by providing fault prediction and fault detection and repairing before a fault occurs, (2) reduce the manual costs and / or undetected faults associated with a maintenance schedule by providing real-time fault detection, (3) reduce human error in sensor maintenance and diagnosis by relying on data, (4) conclusions drawn based on existing evidence, and (5) identify fault tolerance actions / opportunities, such as relying on data collected by correlated sensors (or a “digital twin model” of the sensor of interest) until the sensor of interest is repaired or replaced, based on the fault tolerance identification operation provided herein. The data-driven approach in some aspects may rely on both current and historical data from the sensor of interest and a set of correlated sensors.

[0030] The system, apparatus, and / or method may provide automated root cause analysis. Root cause analysis of failures may be performed manually based on domain knowledge and data visualization, which can be subjective, time-consuming, and error-prone. In some cases, the root cause may be associated with raw sensor data that is not addressed by the domain knowledge or data visualization used in manual root cause analysis. The system, apparatus, and / or method may provide an automated root cause based on a standardized approach to identifying the root cause of predicted failures.

[0031] In some aspects, sensor data (e.g., IoT sensor data, vibration data) may be high-frequency data (e.g., 1000 Hz to 3000 Hz). High-frequency data poses challenges for constructing solutions to failure prediction problems in some aspects. For example, high-frequency data may be associated with high levels of noise or long or resource-consuming analysis (e.g., computation) times. The sampling frequency or aggregation window may require optimization to accurately predict one of a short-term or long-term failure. Accordingly, the system, apparatus, and / or method may provide a window optimization operation to identify an optimized window and / or aggregation statistics for failure prediction.

[0032] In some aspects, physical sensor data may not be able to capture all signals that may be useful for monitoring a system due to the harsh environment for sensor installation, the cost of the sensors, and / or the functionality of the sensors. As a result, the collected data may not be sufficient to monitor the health of the system and capture potential risks and failures. In some aspects, the inability to capture all of these potentially useful signals can pose challenges to building a fault prediction solution. Accordingly, the present system, apparatus, and / or method may enhance the physical sensor data to capture the signals necessary to assist in system monitoring and building a fault prediction solution. For example, the physical sensor data may be processed by a set of physics-based models to generate virtual sensor data.

[0033] FIG. 1 is a diagram 100 showing a solution architecture for fault detection, fault prediction, fault repair, and fault tolerance. The solution architecture may include a sensor data module 110. The sensor data module 110 may incorporate a set of physical sensors 110a and a set of virtual sensors 110b. The physical sensors 110a may include any one or more of a temperature sensor, a pressure sensor, a vibration sensor, an acoustic sensor, a motion sensor, a level sensor, an image sensor, a proximity sensor, a water quality sensor, a chemical sensor, a gas sensor, a smoke sensor, an IR sensor, an acceleration sensor, a gyro sensor, a humidity sensor, a light sensor, and / or a LIDAR sensor. Data from the physical sensors 110a may be provided to a set of physics-based models to generate data associated with the virtual sensors 110b.

[0034] In some embodiments, the physical sensor 110a may be installed on a target asset (e.g., an asset within an OT system) and may be used to collect data for monitoring the health and performance of the asset. Different types of sensors are designed to collect different types of data across different industries, different assets, and different tasks. Different sensors may be included in the physical sensor 110a for different purposes, but the present disclosure generally discusses them because the method can be applied to a wide range of sensors and data types. The virtual sensor 110b may, in some embodiments, be associated with output variables from a set of physical-based models or a set of digital twin models, which can complement and / or verify data from the physical sensors and thus help monitor and maintain system health. In the case of "complementation", if physical sensor data is not available or is insufficient, virtual sensor data from the digital twin model can function as a "substitute" for the physical sensor. In the case of "verification", assuming that the physical sensor also collects data as an output of the digital twin model, the virtual sensor data can function as an "expected" value, while the value from the physical sensor can function as an "observed" value, and thus the variance or difference between them can be used as a signal for detecting abnormal behavior in the system.

[0035] In some embodiments, the data collected by one sensor S1 may be closely related to the data collected by another sensor S2. In this case, S1 can be a substitute for S2, and vice versa. For example, the wind turbine shaft torque can be approximately represented by the amount of vibration generated by the generator, and vice versa. Such a substitution relationship can be obtained based on domain knowledge and / or data analysis (such as correlation analysis). Substitute sensors enable fault tolerance, and if one sensor fails, the other sensor may be used as a substitute to construct a solution. Additionally, a failed sensor may be recognized if such a substitution relationship does not hold.

[0036] Sensor data from the sensor data module 110 may be provided to the critical sensor identification module 120. In some aspects, multiple sensors may be installed on assets within an industrial system to monitor the health of the system, but only some of the sensors are useful for deriving insights, making decisions, and / or constructing solutions for downstream tasks. Such sensors are essential for maintaining a healthy, reliable, and continuously operating industrial system, and it is necessary to keep these sensors functioning as expected. Such sensors may be referred to as critical sensors, and special attention may be paid to these critical sensors beyond what is paid to other non-critical sensors within the system. There are several approaches that may be employed by the critical sensor identification module 120 to identify critical sensors.

[0037] In some aspects, a domain knowledge-based approach may be used. For example, an operator and / or technician may often possess domain knowledge that enables them to identify which sensors are useful and essential for monitoring the health of the system. In some aspects, the operator and / or technician may provide an input to the critical sensor identification module 120 to identify critical sensors. For example, that domain knowledge may be used to identify a list of critical sensors ranked by their importance.

[0038] A data-driven approach may be used to identify one or more critical sensors. For example, one or more variables used to indicate the health of the system may be identified, and data analysis may be performed on historical data from multiple sensors associated with the system to identify which sensors are closely related to the health indicator variables. One approach is to calculate the correlation coefficient between the data of each sensor and the health indicator variables. As a result, a list of sensors ranked by that correlation coefficient can be obtained.

[0039] Sensor data may be used to construct solutions for several downstream tasks including, but not limited to, fault prediction, anomaly detection, remaining useful life, yield optimization, etc. Since the solutions for downstream tasks are constructed based on data from multiple sensors, the importance of sensors for such solutions can be identified through one or more downstream task-based approaches (e.g., feature selection techniques and / or explainable AI techniques). Based on the model constructed based on the downstream tasks, the system and / or method may calculate a value (e.g., a value related to the explanatory impact of sensor data on the downstream tasks) that reflects the feature importance of each sensor. One or more downstream task-based approaches may, in some aspects, provide a list of sensors ranked by their importance.

[0040] The above approaches, such as domain knowledge-based approaches, data-driven approaches, and downstream task-based approaches, may be used independently to identify critical sensors. In some aspects, different approaches may be combined into one approach by merging the ordered lists of sensors generated by the different approaches. For example, one possible approach is to calculate the average ranking of each sensor based on its ranking in the three lists and then sort the sensors based on the average ranking. When calculating the average ranking, a weighted average can be used by first assigning weights to each approach and using the weighted ranks to calculate the average rank.

[0041] The fault tolerance identification module 130 may be used to identify one or more sensors that can provide alternative sensor data for critical sensors. For example, when sensor S1 is provided, a set of one or more sensors that can function as an alternative to sensor S1 may be identified. Based on the set of one or more sensors identified by the fault tolerance identification module 130 and the predicted or detected fault of sensor S1, for a certain period until at least S1 is repaired or replaced, a set of one or more "substitute" sensors may be used in place of sensor S1.

[0042] Figure 2A is a diagram 200 showing steps used to identify sensors that enable fault tolerance. At 210, the method may obtain data for all sensors and take the sensor data values in time series as a vector. The sensors may, in some embodiments, be physical sensors and / or virtual sensors from a digital twin model. At 220, the method may calculate a similarity score between two vectors for each pair of sensors. To perform the comparison, in some embodiments, the data from the pair of sensors is normalized so that they can be compared. For example, the data from the pair of sensors may have been collected during different periods or at different frequencies initially, and the method may sample the data from different sensors to make the time window and data frequency the same. Once the data is normalized, the method may select one or more similarity metrics. The one or more similarity metrics may include, but are not limited to, the correlation coefficient, cosine similarity, Hamming distance, Euclidean distance, Manhattan distance, and / or Minkowski distance. Based on the one or more selected similarity metrics, the method may measure the similarity between the two vectors. Calculating the similarity score between two vectors for a pair of sensors may include one of several approaches. In theory, the method may calculate the similarity score for every possible pair of sensors (e.g., physical sensors and / or virtual sensors), but in some embodiments, the similarity score may be calculated for each of the critical sensors and the other sensors (including both critical and non-critical sensors).

[0043] The similarity score calculated at 220 may be compared with a threshold similarity score at 230 to determine whether the two sensors are correlated and / or related. If the calculated similarity score exceeds the threshold, the similarity may be verified based on the domain knowledge by the operator and / or engineer. The similarity identified at 230 may be referred to as macro-similarity based on comparison of a larger dataset (e.g., data collected over a day, a week, etc.) than that used to identify micro-similarity as will be described below with respect to FIGS. 3A and 3B. As described above, macro-similarity may be used to identify a set of “surrogate” sensors for critical sensors or a set of related and / or correlated sensors for repair or other downstream tasks.

[0044] FIG. 2B is FIG. 205 for bootstrapping the macro-similarity score. If two data vectors from two sensors exceed a threshold length, the similarity calculation can require a lot of resources and an extremely long time to complete. FIG. 205 shows a workflow regarding how to use bootstrapping techniques to calculate the similarity score. FIG. 205 shows that the method may include obtaining data at 240 as described above with respect to step 210 of FIG. 200.

[0045] After obtaining the data, the method may identify whether to analyze the complete dataset or a reduced (e.g., bootstrapped) dataset (not shown). The reduction (e.g., bootstrapping) technique may include sampling corresponding data from each of the two vectors at 250. For example, the method may sample from the two vectors by replacement at a predetermined sampling rate, e.g., 0.01, at 250 and use it to calculate and / or compute the similarity score at 260. Calculating the similarity score at 260 is similar to calculating the similarity described with respect to 220 of FIG. 200 and is performed only over a shorter time window than normal.

[0046] After calculating the similarity score for the current sample, the method may proceed to determine, at 270, whether the threshold number of iterations has been met (each iteration is associated with a sampling-based similarity score. The threshold number of iterations may be set prior to analysis and may be selected to be large enough to ensure that the calculated value is the reported value. If the threshold number of iterations has not been met, the method may return to step 250. Thus, the method may repeat the process of obtaining a plurality of similarity scores multiple times. The method may then aggregate, at 280, the similarity scores from the multiple runs and use the aggregated value as the final similarity score. The aggregation function used at 280 may include, but is not limited to, mean, weighted mean, maximum, minimum, median, weighted median, etc. Based on the aggregated similarity score, at 280, the aggregated similarity score can then be compared to a predetermined similarity score threshold to determine whether the two vectors are similar.

[0047] Figure 3A is a diagram showing the calculation of the micro similarity score. In some embodiments, as described with respect to steps 220 and 280, instead of calculating one similarity score (or an aggregated similarity score), the method may calculate a series of similarity scores based on data within a time window (or time segment). Figure 3A shows a workflow 300 regarding how the micro similarity calculation functions. Regarding generating the macro similarity score in Figure 2A, the method may first, at 310, obtain data for pairs of sensors during the same (or overlapping) time window. At 220, the method may identify the strategy used to define the time window used when calculating the micro similarity score. The time window is, in some embodiments, one of a rolling window or an adjacent window. The time window may also be event-dependent. For example, holiday seasons, business hours of the day, weekdays, weekends, etc. may be used to identify the time window.

[0048] For each time window, the method may calculate a similarity score at 330, resulting in a series of similarity scores for each pair of sensors. The method may then, at 340, obtain the distribution of the similarity scores based on those values and frequencies and analyze the distribution of the similarity scores. To determine whether two sensors are similar, the method may perform a statistical significance test to determine whether a predetermined similarity score threshold is significantly different from the distribution of the similarity scores. For example, the method may use a one-sample one-sided t-test (or other appropriate statistical analysis) to determine whether the similarity score threshold is significantly lower than the similarity scores. The method may first calculate a statistical value based on the data of the similarity score threshold relative to the distribution of the similarity scores. Then, based on the significance level, the method may determine whether the similarity score threshold is significantly lower than the similarity scores. In this case, a one-sided test is used, i.e., focusing on the left critical region in the distribution of the similarity scores. Micro similarity provides a more detailed view of the similarity scores, thus providing more information to represent the similarity between two sensors and being accurate.

[0049] Figure 3B is Figure 305 showing a method of bootstrapping based on micro similarity. Similar to the relationship between Figures 2A and 2B, Figure 3B shows that the first two steps of the method, namely obtaining data at 350 and specifying a strategy to define a time window at 355, are equivalent to steps 310 and 320 respectively. In the micro similarity approach, if there are too many time windows, it may consume a lot of resources for calculation and take too much time to execute. Therefore, the method shown in Figure 305 may use a bootstrapping technique to solve such problems. When the method specifies a window generation strategy at 355 and defines all time windows, at 360, the time windows can be sampled using a bootstrapping technique by substitution with a predefined sampling rate, for example 0.01. Then, the method may apply the micro similarity approach at 365 to calculate a series of similarity scores and the distribution of similarity scores. The method may compare a similarity score threshold with the distribution of similarity scores based on a statistical significance test at 370, and the result of the current execution is recorded. At 375, the method determines whether additional executions should be performed. If so, the method may return to 360 and perform another random sampling of the time windows defined at 355. The sampling execution may continue several executions of bootstrapping sampling and application of the micro similarity approach until a predetermined number of executions is reached. The results from the predetermined number of executions may be aggregated at 380 to obtain a final result. Since the result from each execution is a binary value indicating whether the similarity score significantly falls below the similarity score (meeting the threshold criteria for identifying similarity through the identity score), in some aspects, a "majority vote" technique is used to check which binary value dominates the result and use it as the final result. In other aspects, when the result from each execution is represented by a numerical score indicating statistical significance, the average or weighted average technique can be used to calculate the average statistical significance value as the final result.Finally, in some embodiments, determining whether two sensors are similar may include that if the calculated similarity score exceeds a threshold, the similarity may be verified based on domain knowledge by an operator and / or engineer. Generally speaking, the approaches of calculating bootstrap similarity, micro similarity, and bootstrap micro similarity each transform the original calculation for a large vector into multiple calculations for small vectors, which reduces the hardware requirements. As a result, the analysis may be able to be performed on edge devices (e.g., devices that may have limited hardware resources) by these approaches.

[0050] In some embodiments, when sensor S1 is provided, there may be no single sensor that can function as a substitute for sensor S1, and the method may select an entire group of sensors that can be used as a substitute for sensor S1. One approach is to use sensor S1 as the target and the remaining sensors as features to build a machine learning model. If the model performance metric exceeds a certain threshold, it can be said that sensor S1 can be substituted by a set of one or more correlated (or related) sensors. To identify the substitute sensors, the method can select important features from the model and use the corresponding sensors as substitute sensors for sensor S1. Feature selection can be performed based on some feature selection techniques including, but not limited to, forward selection, backward selection, model-based feature selection, etc. Domain knowledge may be incorporated in some embodiments to improve feature selection. The group of sensors used as a substitute for the target sensor (in this case S1) is called a cohort sensor, a related sensor, or a correlated sensor. In addition to using a group of cohort sensors as a substitute for the target sensor, the output from a machine learning model (or a set of physics-based models associated with virtual sensors) can also be used as a substitute for the target sensor.

[0051] In addition to identifying the similarity between sensors based on sensor data, the methods described herein may incorporate some domain knowledge, if available. For example, some exemplary sensors that may have a high degree of similarity are: a) physical sensors for input variables and input variables within a designed motion profile; b) physical sensors for output variables and output variables from a digital twin model (which may be based on input variables from either the physical sensors for the designed motion profile or the input variables); and / or c) output variables from different versions of the digital twin model (which may be based on input variables from either the physical sensors for the designed motion profile or the input variables).

[0052] The methods described with respect to FIGS. 2A - 3B may be related to and / or performed by the critical sensor identification module 120 and / or the fault tolerance identification module 130. Based on the output of the fault tolerance identification module 130 and in some aspects the critical sensor identification module 120, the fault detection module 140 may perform a fault detection operation. For example, after a critical sensor is identified by the critical sensor identification module 120 and a similar sensor is identified (if possible) for the critical sensor by the fault tolerance identification module 130, one or more data-driven approaches are used to detect faults in the critical sensor. The data may be physical sensor data and / or virtual sensor data from the digital twin. This approach may, in some aspects, involve one or more machine learning models including a univariate anomaly detection model, a bivariate anomaly detection model, and / or a multivariate anomaly detection model.

[0053] For the sensor of interest, the univariate anomaly analysis may include running an anomaly detection model against the sensor's data. Anomaly in the temporal sequence of the data may indicate either a faulty sensor or an operational anomaly. Thus, the second anomaly analysis is, in some aspects, performed to distinguish between a faulty sensor and an operational anomaly. For example, FIG. 4A shows a method 400 of bivariate analysis that identifies whether related and / or corresponding sensors have experienced (or are experiencing) similar problems indicating an operational anomaly, or have not experienced (or are not experiencing) similar problems indicating that the sensor is faulty. For example, at 410, for the sensor of interest, the method first identifies similar sensors based on the output from the fault tolerance identification module 130. For cohort sensors, the method may use the output from a machine learning model based on the group of cohort sensors as the similar sensors.

[0054] At 420, the method may then run a micro similarity algorithm against the historical data from the sensor of interest and the similar sensors, as described above with respect to FIGS. 3A and 3B above, to calculate a series of similarity scores. The method may also obtain the distribution f of the similarity scores. After running the micro similarity algorithm against the historical data, at 430, the method runs the micro similarity against the new data from the sensor of interest and the similar sensors and obtains the current similarity scores.

[0055] Finally, at 440, the method may check whether the similarity scores based on the historical data and the similarity scores based on the current data differ to the extent of indicating a faulty sensor. For example, an anomaly detection model may be run against the series of similarity scores to detect such differences or anomalies, and / or the method may perform a statistical significance test of the similarity scores against the distribution of the similarity scores. A one-sample t-test may be performed by selecting a significance level, e.g., 0.01. Anomaly detected by the bivariate anomaly detection model usually indicates that there is a fault in the sensor of interest or the similar sensors.

[0056] In some embodiments, data from multiple sensors may be used to build a multivariate anomaly detection model. Such anomalies typically indicate system operation anomalies, assuming that it is unlikely for multiple sensors to fail simultaneously. Among the three approaches described above, the univariate anomaly detection model may not be able to distinguish anomalies caused by sensor failures or system operation failures, the bivariate anomaly detection model may not be able to identify which of the two sensors has a failure, and the multivariate anomaly detection model only detects system operation anomalies. To address these limitations, some approaches use an ensemble or combination of the above approaches to identify which sensors have failures. We introduce two approaches for this purpose. Each approach can be executed independently to detect sensor failures.

[0057] Figure 4B is a diagram 405 showing a first ensemble approach for sensor fault detection. In the first ensemble approach shown in diagram 405, the outputs from a univariate anomaly detection model executed at 450 and a bivariate anomaly detection model executed at 470 may be used to detect a fault in the sensor. This approach utilizes the presence of related or corresponding sensors with respect to the target sensor. As described above, similar (e.g., related or correlated) sensors may be detected by the fault tolerance identification module 130. For example, referring to Figure 4B, the first ensemble approach may include, at 450, executing a univariate anomaly detection model on a vector of sensor data for a particular target sensor (e.g., a critical sensor). Based on the univariate anomaly detection model executed at 450, the method may, at 460, identify whether an anomaly has been detected. If no anomaly is detected, the sensor may be identified at 460 as not being faulty at 490B. However, if an anomaly is detected by the univariate anomaly detection model executed at 450, the method may, at 470, execute a bivariate sensor anomaly detection model on the vectors of the target sensor and related / correlated / similar sensors. Based on the bivariate anomaly detection model executed at 470, the method may, at 480, identify whether an anomaly (e.g., an anomaly between the outputs of the target sensor and related / correlated / similar sensors) has been detected. If no anomaly is detected (e.g., the related / correlated / similar sensors generate measurement values / vectors similar to the measurement values / vectors generated by the sensor of the target sensor), the target sensor may be identified at 480 as not being faulty at 490B. However, if an anomaly is detected based on the bivariate anomaly detection model executed at 470, there is evidence that the target sensor (or critical sensor) data does not match the sensor data collected by the related / correlated / similar sensors, and since the univariate anomaly detection model identifies an anomaly in the target sensor and it can be concluded that the target sensor is faulty, the target sensor may be identified at 480 as being faulty at 490A.

[0058] FIG. 5 is a diagram 500 showing a second ensemble approach for sensor fault detection. In the second ensemble approach shown in FIG. 500, the outputs from a univariate anomaly detection model executed at 550 and a multivariate anomaly detection model executed at 570 may be used to detect faults in the sensors. The second ensemble approach may include, at 550, executing a univariate anomaly detection model against a vector of sensor data for a particular sensor of interest (e.g., a critical sensor). Based on the univariate anomaly detection model executed at 550, the method may, at 560, identify whether an anomaly has been detected. If no anomaly has been detected, the sensor may be identified at 560 as not being faulty at 590B. However, if an anomaly is detected by the univariate anomaly detection model executed at 550, the method may execute a multivariate sensor anomaly detection model against the vectors of all (critical) sensors, including the sensor of interest. Based on the multivariate anomaly detection model executed at 570, the method may, at 580, identify whether an anomaly has been detected. If an anomaly has been detected, the sensor may be identified at 580 as not being faulty at 590B (e.g., if the multivariate anomaly detection model based on all (critical) sensors detects an anomaly, it is likely to be a system / operation anomaly and not a sensor fault). However, if no anomaly is detected based on the multivariate anomaly detection model executed at 570, the multivariate anomaly detection model does not identify any system / operation anomalies and thus it is likely to be a sensor fault, and the sensor may be identified at 580 as being faulty at 590A.

[0059] In some embodiments, additional considerations may be used to identify a faulty sensor. For example, if a sensor is unable to generate any data readout values, the sensor may be identified as faulty. In some embodiments, the first and second ensemble approaches may be executed simultaneously to detect a faulty sensor. If both approaches detect a fault in the sensor, the sensor may be identified as faulty. If both approaches are unable to detect a fault in the sensor, the sensor may be identified as not faulty. If one approach detects a fault in the sensor and the other does not, either a faulty sensor or a non-faulty sensor is output depending on the risk-versus-cost trade-off.

[0060] When building an anomaly detection model, time series data may, in some embodiments, be preprocessed before applying the model to the preprocessed data. Some preprocessing techniques may include, but are not limited to, differencing, moving average, moving variance, window-based features, etc. The approach can be applied to both analog and digital sensors. For digital sensors, the data can first be preprocessed by moving average and / or moving variance, after which the data becomes continuous values.

[0061] Anomaly detection may, in some embodiments, be a distribution-based method. For example, a moving variance based on sensor data may first be calculated, and then the distribution of the moving variance may be calculated or identified. Based on the distribution of the moving variance, anomaly detection (e.g., performed by an anomaly detection module) may identify outliers / anomalies based on a predetermined threshold (e.g., outside the 99% range). Such outliers / anomalies may, in some embodiments, be identified as corresponding to a faulty sensor. The assumption here is that if the sensor data remains at the same value for some time (i.e., the moving variance is close to 0), a deviation from that value corresponds to a fault in the sensor.

[0062] By means of a fault detection technique, a fault may be detected when the fault occurs. While the fault sensor is being repaired and / or replaced, the underlying system may be left unmonitored due to the sensor downtime. To avoid leaving the underlying system unmonitored during maintenance / repair, in some aspects, a fault prediction module is provided to predict a sensor fault in advance to avoid the sensor fault or to enable repair without downtime where the system is unmonitored. FIG. 6 is a diagram 600 showing an exemplary fault prediction module 650. In some aspects, the fault prediction module 650 corresponds to the fault prediction module 150 of FIG. 1. The fault prediction module 650 may execute a set of anomaly detection models 651a and / or 652a to obtain a corresponding set of anomaly scores. The set of anomaly detection models may include one or more of a univariate anomaly detection model for the data of each sensor, a bivariate anomaly detection model for each pair of identified related / correlated / similar sensors, or a multivariate anomaly detection model, as described above, to generate a set of anomaly scores.

[0063] The fault prediction module 650 may identify and / or prepare (e.g., in the feature module 651) features associated with a set of sensors (e.g., associated with sensor data 651b). For example, for each sensor, the fault prediction module 650 may obtain the following data: (1) data from the sensor and similar sensors (if available), (2) an anomaly score from a univariate anomaly detection model, (3) an anomaly score from a bivariate anomaly detection model, and (4) an anomaly score from a multivariate anomaly detection model. Additionally, for each sensor, the fault prediction module 650 may prepare (e.g., in the target module 652) a target, and preparing the target may include obtaining, for a related lead time, one or more of (1) an anomaly score from a univariate anomaly detection model, (2) an anomaly score from a bivariate anomaly detection model, and / or (3) an anomaly score from a multivariate anomaly detection model, where the lead time may be a predetermined value indicating how far in advance an anomaly is predicted. Based on the obtained data, the fault prediction module 650 may construct a sequence prediction model (e.g., the fault prediction model 653). The sequence prediction model may, in some embodiments, be a deep learning recurrent neural network (RNN) used for sequence prediction. The deep learning RNN may be one of a long short-term memory (LSTM) model or a gated recurrent unit (GRU) model. The deep learning RNN model may, in some embodiments, enable multiple targets at once such that the output from each prediction includes three anomaly scores, a univariate, bivariate, and multivariate anomaly score (e.g., associated with the predicted anomaly score 655). Finally, the ensemble approach described above with respect to FIGS. 4B and 5 may be applied by the ensemble prediction anomaly score module 657 to predict whether a sensor is faulty.

[0064] The fault repair and fault tolerance module 160 may receive information from one or more of the critical sensor identification module 120, the fault tolerance identification module 130, and / or the fault prediction module 150. When a fault is predicted, the fault repair and fault tolerance module 160 may identify the root cause of the fault using explainable AI techniques (such as ELI5 and Shapley Additive Explanation (SHAP)). In this case, the fault repair and fault tolerance module 160 may identify the data points (or critical sensors) associated with the predicted fault and / or the root cause. The post-processing of the predicted anomaly score and the root cause may verify the fault with or without applying domain knowledge. If the fault is valid, the fault repair and fault tolerance module 160 may identify one or more repairs that can be performed by an operator or technician (including identifying a "substitute" sensor). For example, identifying one or more repairs may include checking whether the faulty sensor has a set of related / correlated / similar sensors to enable fault tolerance. If a set of related / correlated / similar sensors exists, the one or more identified repairs may include using the set of related / correlated / similar sensors for downstream tasks and replacing the faulty sensor. If a set of related / correlated / similar sensors does not exist and / or cannot be identified, the one or more identified repairs may include immediately replacing the faulty sensor and adding at least one additional sensor to provide redundancy in the future. The one or more identified repairs may include location-based faulty sensor repair such that if sensors of the same type exist continuously (e.g., upstream or downstream of the faulty sensor), the upstream and / or downstream sensors can be used to complement the faulty sensor value. The one or more identified repairs may also include time-based faulty sensor repair such that if for some reason the sensor generates faulty values only over a specific period, the data before and after the fault can be used to complement the sensor value during the fault period.

[0065] In some embodiments, fault tolerance may be systematically introduced during design time and / or operation time. For example, a digital twin model may be constructed to output virtual sensor data as a substitute for fault tolerance for physical sensor data. In some embodiments, virtual sensors from the digital twin model may complement and assist in validating physical sensors. Two or more versions of the digital twin model may be constructed in some embodiments to enable more fault tolerance and data validation. In some embodiments, the introduction of fault tolerance may include identifying critical sensors and introducing at least one additional sensor for fault tolerance if similar sensors are not available. For example, in a service-oriented architecture (SOA), it may be desirable to ensure fault tolerance for critical sensors.

[0066] FIG. 7 is a flowchart 700 of a method for detecting and repairing faults in sensors associated with a system. The method may be performed by one or more components of a system such as the solution architecture shown in FIG. 100 or an individual or combined solution architecture, such as sensor data module 110, critical sensor identification module 120, fault tolerance identification module 130, fault detection module 140, fault prediction module 150 (or 650), and / or fault repair and fault tolerance module 160. At 710, the method may receive sensor data from a plurality of associated sensors. In some embodiments, the plurality of associated sensors are sensors that monitor the same system. The plurality of associated sensors may include, in some embodiments, at least one virtual sensor derived from a set of physical sensors based on physical sensors installed in the system or a physical-based model. For example, the sensor data may be received by sensor data module 110 from a plurality of associated sensors, or may be received from sensor data module 110 by another module of the solution architecture.

[0067] In 720, the method may identify a set of correlated sensors among a plurality of related sensors for a first sensor among the plurality of related sensors. The first sensor, in some aspects, is a first critical sensor, and a critical sensor captures critical data for monitoring the health of the underlying system and may be used for at least one of identifying a repair strategy, deriving business insights, or constructing a solution for problems related to a set of downstream tasks. Downstream tasks may, in some aspects, include one or more of anomaly detection, fault prediction, remaining useful life prediction. In some aspects, the set of correlated sensors includes a set of sensors having outputs that correlate with the output of the first sensor. The set of correlated sensors, in some aspects, includes a plurality of sensors among the plurality of related sensors. For example, in some aspects, the correlation may be computed between critical sensor data and the first principal component of the plurality of sensor data (or some other function of the plurality of sensor data even if the individual sensor data do not correlate at a threshold level). As shown by the expansion view of 720, for identifying a set of correlated sensors among a plurality of related sensors in 720, in some aspects, the method may further compute, in 720A, a similarity score between sensor data from the first sensor and sensor data from sensors among the plurality of related sensors, and may identify, in 720B, the set of correlated sensors based on the similarity score between sensor data from the first sensor and sensor data from the set of correlated sensors exceeding a threshold similarity score.

[0068] FIG. 8 is FIG. 800 which further expands a diagram of sub-operations performed to identify a set of correlation sensors among a plurality of related sensors at 720 in some embodiments. Elements 820, 820A, and 820B of FIG. 8 correspond to elements 720, 720A, and 720B of FIG. 7 in some embodiments and may further identify sub-operations associated with calculating a similarity score at 720A / 820A. At 820A-1, the method may calculate a macro similarity score based on a full set of time-series sensor data from a first sensor during a first period and a full set of time-series data from each sensor among a plurality of related sensors during the first period. Additionally, at 820A-2, the method may calculate a plurality of micro similarity scores based on a plurality of subsets of the full set of time-series data from the first sensor during the first period and a plurality of subsets of the full set of time-series data from each sensor among a plurality of related sensors during the first period. As described above with respect to FIGS. 2A-3B, the calculation may be "standard" as in FIGS. 2A and 3A, or "bootstrapped" as in FIGS. 2B and 3B.

[0069] At 720B / 820B, the method may identify a set of correlation sensors based on a similarity score between sensor data from the first sensor and sensor data from the set of correlation sensors exceeding a threshold similarity score. The set of correlation sensors may then be identified as a set of one or more "substitute" sensors for a critical sensor in the event of a sensor failure. For example, 720 may be performed by a fault tolerance identification module 130.

[0070] In 730, the method may detect a fault in the first sensor based on at least one of the sensor data received from the sensor, the sensor data received from a set of correlated sensors, and the sensor data received from other sensors. The fault detection in 730 may be performed by the fault detection module 140 or the fault prediction module 150. In some aspects, in 730, detecting a fault in a sensor based on at least one of the sensor data received from the sensor, the sensor data received from a set of correlated sensors, and the sensor data received from other sensors includes, in 730A, using one or more of the following models to detect a fault in the sensor: (1) a univariate anomaly detection model based on the sensor data received from the first sensor, (2) a bivariate anomaly detection model based on the sensor data received from the first sensor and the sensor data received from a set of correlated sensors, and / or (3) a multivariate anomaly detection model based on the sensor data received from the first sensor, the sensor data received from a set of correlated sensors, and the sensor data received from all sensors including other sensors in the system. In some aspects, the method may further include, in 730B, using one or more of the following ensemble models to detect a fault in the sensor in 730: (1) an ensemble anomaly detection model based on the univariate anomaly detection model and the bivariate anomaly detection model, or (2) an ensemble anomaly detection model based on the univariate anomaly detection model and the multivariate anomaly detection model.

[0071] In an aspect of detecting real-time faults in the first sensor, the univariate (anomaly) score, the bivariate (anomaly) score, and / or the multivariate (anomaly) score may be calculated in 730A based on a univariate anomaly detection model, a bivariate anomaly detection model, and / or a multivariate anomaly detection model, respectively. In some aspects, the ensemble (anomaly) score may be calculated in 730B based on one or both of (1) an ensemble anomaly detection model based on the univariate anomaly score and the bivariate anomaly score, and / or (2) an ensemble anomaly detection model based on the univariate anomaly score and the multivariate anomaly score. In some aspects, detecting a fault in the first sensor at 730 includes detecting a predicted fault of the first sensor. Detecting a predicted fault of the first sensor may, in some aspects, be based on a sequence prediction model based on a deep learning RNN as discussed above. In some aspects, the sequence prediction model may be one of an LSTM or GRU model. As part of detecting a predicted fault in the first sensor at 730, the method may calculate (or predict) a plurality of predicted anomaly scores including at least two of a univariate anomaly score, a bivariate anomaly score, or a multivariate anomaly score based on a prediction model (e.g., the sequence prediction model described above) in 730A. The method may further calculate a predicted fault score indicating the likelihood of a sensor fault by generating an ensemble of anomaly scores in 730B based on the plurality of anomaly scores calculated in 730A and one or both of (1) an ensemble anomaly detection (prediction) model based on the univariate anomaly score and the bivariate anomaly score, and / or (2) an ensemble anomaly detection (prediction) model based on the univariate anomaly score and the multivariate anomaly score. For example, 730, 730A, and 730B may be executed by the fault prediction module 150.

[0072] Finally, at 740, the method may implement a repair strategy (e.g., start, propose to the user to implement, etc.) based on the failure of the first sensor detected at 730. For example, 740 may be executed by the fault repair and fault tolerance module 160. In some aspects, the repair strategy includes replacing the sensor data from the first sensor with sensor data from one or more sensors within the set of correlated sensors. In some aspects, the repair strategy is based on a root cause analysis of the detected failure based on one or more explainable artificial intelligence (AI) techniques (e.g., ELI5 or SHAP). Given the prediction results from the machine learning model, the explainable AI technique can discover the contribution of each feature (used in the machine learning model) to the prediction results. The contribution is measured by a weight value that can be used by the operator and / or technician to elucidate the root cause of the prediction results. The sensor data from one or more sensors within the set of correlated sensors is identified based on the following criteria: (1) the calculated similarity score between the sensor data from the first sensor and the sensor data from the sensors within the set of related sensors, (2) the geolocation of the first sensor and the sensors within the set of related sensors, and / or (3) one or more of the time series of the first sensor and the sensors within the set of related sensors.

[0073] As described above, the present disclosure introduces several data-driven approaches for automatically detecting, predicting, and repairing sensor failures in real time. Fault detection, prediction, repair, and resilience are provided as needed (e.g., just-in-time models), avoiding unnecessary inspections and avoiding unnecessary monitoring or maintenance of non-critical sensors while being applied in real time to a set of critical sensors. Both physical sensors and / or virtual sensors (from digital twin models) are incorporated into this solution framework, and virtual sensors, in some aspects, provide fault tolerance to physical sensors. The present disclosure enables an operator or technician to distinguish faults in sensor-to-system components. Additionally, fault tolerance from already installed similar sensors is identified and utilized to reduce the cost of introducing new sensors to enable fault tolerance or the cost associated with failure of critical sensors. Further, sensor failures can be predicted, thereby enabling repair to be performed or scheduled at a more convenient or less costly time (e.g., before incurring costs associated with sensor failure). The analysis may also be used to proactively introduce systematic fault tolerance presented before any faults are detected.

[0074] FIG. 9 shows an exemplary computing environment having an exemplary computer device suitable for use in some exemplary implementations. The computer device 905 within the computing environment 900 can include one or more processing units, cores, or processors 910, memory 915 (e.g., RAM, ROM, and / or the like), internal storage 920 (e.g., magnetic, optical, solid state storage, and / or organic), and / or an IO interface 925, any of which can be connected on a communication mechanism or bus 930 for communicating information or can be incorporated within the computer device 905. The IO interface 925 is also configured to receive images from a camera or provide images to a projector or display, depending on the desired implementation.

[0075] The computer device 905 can be communicably connected to an input / user interface 935 and an output device / interface 940. One or both of the input / user interface 935 and the output device / interface 940 can be a wired or wireless interface and can be detachable. The input / user interface 935 can include any physical or virtual device, component, sensor, or interface (e.g., buttons, touch screen interface, keyboard, pointing / cursor control, microphone, camera, Braille, motion sensor, accelerometer, optical reader, and / or the like) used to provide an input. The output device / interface 940 can include a display, television, monitor, printer, speaker, Braille, or the like. In some exemplary implementations, the input / user interface 935 and the output device / interface 940 can be incorporated with or physically connected to the computer device 905. In other exemplary implementations, other computer devices can function as or provide the functions of the input / user interface 935 and the output device / interface 940 for the computer device 905.

[0076] Examples of the computer device 905 can include, but are not limited to, devices with high mobility (e.g., smartphones, devices within vehicles or other machines, devices carried by humans and animals, and the like), mobile devices (e.g., tablets, notebooks, laptops, personal computers, portable televisions, radios, and the like), and devices not designed for mobility (e.g., desktop computers, other computers, information kiosks, televisions, radios, and the like having one or more processors incorporated therein and / or connected thereto).

[0077] The computer device 905 can be communicatively connected (e.g., via the IO interface 925) to an external storage 945 and a network 950 for communicating with any number of network-connected components, devices, and systems including one or more computer devices of the same or different configurations. The computer device 905 or any connected computer device can function as, provide the functionality of, or be referred to as a server, client, synserver, general-purpose machine, special-purpose machine, or another label.

[0078] The IO interface 925 can include wired and / or wireless interfaces that use any communication or IO protocol or standard (e.g., Ethernet, 802.11x, Universal Serial Bus, WiMax, modem, cellular network protocol, and the like) for communicating information to and / or from at least all connected components, devices, and networks within the computing environment 900. The network 950 can be any network or combination of networks (e.g., the Internet, local area network, wide area network, telephone network, cellular network, satellite network, and the like).

[0079] The computer device 905 can use and / or communicate using computer-usable or computer-readable media including transient media and non-transient media. Transient media includes transmission media (e.g., metal cables, optical fibers), signals, carrier waves, and the like. Non-transient media includes magnetic media (e.g., disks and tapes), optical media (e.g., CD-ROM, digital video disk, Blu-ray disk), solid-state media (e.g., RAM, ROM, flash memory, solid-state storage), and other non-volatile storage or memory.

[0080] Computer device 905 can be used to implement techniques, methods, applications, processes, or computer-executable instructions in some exemplary operating environments. Computer-executable instructions can be obtained from a temporary medium, stored on a non-temporary medium, and obtained therefrom. The executable instructions can be derived from one or more of any programming, scripting, and machine languages (e.g., C, C++, C#, Java, Visual Basic, Python, Perl, JavaScript, and others).

[0081] Processor 910 can execute under any operating system (OS) (not shown) in a native or virtual environment. Along with the OS and other applications (not shown), one or more applications can be arranged including logic unit 960, application programming interface (API) unit 965, input unit 970, output unit 975, and inter-unit communication mechanism 995 for different units to communicate with each other. The described units and elements can be changed in design, function, configuration, or implementation form and are not limited to the provided description. Processor 910 can be in the form of a hardware processor such as a central processing unit (CPU) or a combination of hardware and software units.

[0082] In some exemplary implementations, when information or execution instructions are received by the API unit 965, they may be transmitted to one or more other units (e.g., the logic unit 960, the input unit 970, the output unit 975). In some cases, the logic unit 960 may be configured to control the information flow between units and the services provided by the API unit 965, the input unit 970, and the output unit 975 in some of the exemplary implementations described above. For example, the flow of one or more processes or implementations may be controlled by the logic unit 960 alone or in cooperation with the API unit 965. The input unit 970 may be configured to obtain inputs for the calculations described in the exemplary implementation, and the output unit 975 may be configured to provide outputs based on the calculations described in the exemplary implementation.

[0083] Processor 910 may be configured to receive sensor data from a plurality of related sensors. Processor 910 may also be configured to identify a set of correlated sensors in the plurality of related sensors for a first sensor among the plurality of related sensors. Processor 910 may be further configured to detect a failure in the first sensor based on at least one of the sensor data received from the sensor, the sensor data received from the set of correlated sensors, and the sensor data received from other sensors. Processor 910 may be further configured to implement a repair strategy based on the detected failure of the first sensor. Processor 910 may be further configured to calculate a similarity score between the sensor data from the first sensor and the sensor data from the sensors in the plurality of related sensors. Processor 910 may also be configured to identify the set of correlated sensors based on the similarity score between the sensor data from the first sensor and the sensor data from the set of correlated sensors exceeding a threshold similarity score. Processor 910 may be configured to calculate a macro similarity score based on the full set of time-series sensor data from the first sensor during a first period and the full set of time-series data from each sensor in the plurality of related sensors during the first period. Processor 910 may also be configured to calculate a plurality of micro similarity scores based on a plurality of subsets of the full set of time-series data from the first sensor during the first period and a plurality of subsets of the full set of time-series data from each sensor in the plurality of related sensors during the first period. Processor 910 may be configured to calculate (or predict) a plurality of anomaly scores including at least two of a univariate anomaly score, a bivariate anomaly score, or a multivariate anomaly score. Processor 910 may also be configured to generate a prediction failure score indicating the likelihood of a sensor failure by generating an ensemble of anomaly scores based on the plurality of anomaly scores.

[0084] Some portions of the detailed description are presented in terms of algorithms and symbolic representations of operations within a computer. These algorithmic descriptions and symbolic representations are the means used by those skilled in the data processing arts to convey the essence of their innovations to others skilled in the art. An algorithm is a defined set of steps that produces a desired end state or result. In one implementation, the steps executed require physical operations of tangible quantities to achieve a specific result.

[0085] Unless otherwise specified, as will be apparent from the description, descriptions that utilize terms such as "processing," "computing," "calculating," "determining," "displaying," etc. throughout the description may include actions and processes of a computer system or other information processing device that manipulate data represented as physical (electronic) quantities within the registers and memories of the computer system, and transform the data into other data similarly represented as physical quantities within the memories or registers of the computer system, or in other information storage, transmission, or display devices.

[0086] The implementation examples may also relate to an apparatus for performing the operations of this specification. This apparatus can be specially constructed for the required purposes, or can include one or more general-purpose computers selectively activated or reconfigured by one or more computer programs. Such computer programs may be stored in a computer-readable medium such as a computer-readable storage medium or a computer-readable signal medium. The computer-readable storage medium may include, but is not limited to, tangible media such as optical disks, magnetic disks, read-only memories, random access memories, solid-state devices and drives, or any other type of tangible or non-transitory medium suitable for storing electronic information. The computer-readable signal medium may include a medium such as a carrier wave. The algorithms and representations presented herein are not inherently related to any particular computer or other apparatus. The computer program may include a pure software implementation that includes instructions for performing the operations of the desired implementation form.

[0087] A variety of general-purpose systems may be used with the programs and modules according to the examples herein, or it may prove convenient to construct more specialized apparatus for performing the desired method steps. Additionally, the implementation examples are not described with respect to any particular programming language. It will be understood that a variety of programming languages may be used to implement the teachings of the implementation examples described herein. Instructions in a programming language may be executed by one or more processing units, such as a central processing unit (CPU), a processor, or a controller.

[0088] As is known in the art, the operations described above can be performed by hardware, software, or some combination of software and hardware. Various aspects of the implementation examples may be implemented using circuitry and logic devices (hardware), while other aspects may be implemented using instructions stored on a machine-readable medium (software), and when such instructions are executed by a processor, cause the processor to execute a method for implementing the implementation forms of the present application. Further, some implementation examples of the present application may be executed by hardware only, while other implementation examples may be executed by software only. Further, the various functions described may be executed within a single unit or distributed among several components in any number of ways. When executed by software, the method may be executed by a processor such as a general-purpose computer based on instructions stored on a computer-readable medium. If desired, the instructions may be stored on the medium in compressed form and / or encrypted form.

[0089] Furthermore, other implementation forms of the present application will become apparent to those skilled in the art by considering this specification and practicing the techniques of the present application. The various aspects and / or components of the described implementation examples may be used alone or in any combination. This specification and the implementation examples are intended to be considered as examples only, and the true scope and spirit of the present application are indicated by the appended claims.

Claims

1. Receiving sensor data from a plurality of associated sensors; For a first sensor among the plurality of associated sensors, identifying a set of correlated sensors among the plurality of associated sensors; Detecting a failure in the first sensor based on at least one of the sensor data received from the sensors, the sensor data received from the set of correlated sensors, and the sensor data received from other sensors; Implementing a repair strategy based on the detected failure of the first sensor; A method comprising the above.

2. Detecting the failure in the first sensor includes detecting a predicted failure of the first sensor, and detecting the predicted failure of the first sensor is based on a sequence prediction model based on a deep learning recurrent neural network, and the sequence prediction model Predicting a plurality of anomaly scores including at least two of a univariate anomaly score, a bivariate anomaly score, or a multivariate anomaly score; Calculating a predicted failure score indicating the likelihood of the sensor failure by generating an ensemble of anomaly scores based on the plurality of anomaly scores. The method according to claim 1, comprising the above.

3. The method according to claim 1, wherein the plurality of associated sensors are sensors that monitor the same system.

4. The method according to claim 1, wherein the plurality of associated sensors include at least one virtual sensor derived from a set of physical sensors based on physical sensors installed in the system or physical-based models.

5. The first sensor is a first critical sensor, and a critical sensor is a sensor that captures critical data for monitoring the health of the underlying system and is used for at least one of identifying the repair strategy, deriving business insights, or constructing solutions for problems related to a set of downstream tasks. The method according to claim 1.

6. The method according to claim 1, wherein the set of correlated sensors includes a set of sensors having outputs correlated with the output of the first sensor.

7. The set of correlation sensors includes a plurality of sensors among the plurality of associated sensors, the output of at least one of the plurality of sensors does not correlate with the output of the first sensor having a threshold correlation, and identifying the set of correlation sensors is based on a function of the outputs of the plurality of sensors that correlate with the first sensor having the threshold correlation, the method according to claim 1.

8. Identifying the set of correlation sensors comprises: calculating a similarity score between the sensor data from the first sensor and the sensor data from sensors in the plurality of associated sensors; and identifying the set of correlation sensors based on the similarity score between the sensor data from the first sensor and the sensor data from the set of correlation sensors exceeding a threshold similarity score. The method according to claim 1, comprising:

9. Calculating the similarity score between the sensor data from the first sensor and the sensor data from sensors in the plurality of associated sensors comprises: calculating a macro similarity score based on a full set of time-series sensor data from the first sensor during a first period and a full set of time-series data from each sensor in the plurality of associated sensors during the first period, or calculating a plurality of micro similarity scores based on a plurality of subsets of the full set of the time-series data from the first sensor during the first period and a plurality of subsets of the full set of the time-series data from each sensor in the plurality of associated sensors during the first period. The method according to claim 8, comprising at least one of:

10. Detecting the fault in the sensor based on at least one of the sensor data received from the sensor, the sensor data received from the set of correlation sensors, and the sensor data received from other sensors includes one or more models for detecting the fault in the sensor, the one or more models being: a univariate anomaly detection model based on the sensor data received from the first sensor; a bivariate anomaly detection model based on the sensor data received from the first sensor and the sensor data received from the set of correlation sensors, or A multivariate anomaly detection model based on the sensor data received from the plurality of related sensors, including the sensor data received from the first sensor, the sensor data received from the set of correlation sensors, and the sensor data received from other sensors The method according to claim 1, comprising:

11. Further comprising using one or more ensemble models to detect the failure in the sensor, wherein the one or more ensemble models are An ensemble anomaly detection model based on the univariate anomaly detection model and the bivariate anomaly detection model, or An ensemble anomaly detection model based on the univariate anomaly detection model and the multivariate anomaly detection model The method according to claim 10, comprising:

12. The method according to claim 1, wherein the repair strategy includes replacing the sensor data from the first sensor with sensor data from one or more sensors within the set of correlation sensors

13. The method according to claim 12, wherein the repair strategy is based on a root cause analysis of the detected failure based on one or more explainable artificial intelligence (AI) techniques

14. The sensor data from the one or more sensors within the set of correlation sensors is A calculated similarity score between the sensor data from the first sensor and the sensor data from the sensors within the set of correlation sensors, The geolocation of the first sensor and the sensors within the set of correlation sensors, or The time series of the first sensor and the sensors within the set of correlation sensors The method according to claim 12, wherein the sensor data is identified based on one or more of the above

15. A device including a memory and At least one processor connected to the memory, wherein based at least in part on information stored in the memory, the at least one processor Receives sensor data from a plurality of related sensors, Identifies a set of correlation sensors in the plurality of related sensors for a first sensor in the plurality of related sensors, Detects a failure in the first sensor based on at least one of the sensor data received from the sensors, the sensor data received from the set of correlation sensors, and the sensor data received from other sensors, and Is configured to implement a repair strategy based on the detected failure of the first sensor Device

16. The at least one processor is configured to, calculate a similarity score between the sensor data from the first sensor and the sensor data from sensors in the plurality of associated sensors, calculate a macro similarity score based on a full set of time-series sensor data from the first sensor during a first period and a full set of time-series data from each sensor in the plurality of associated sensors during the first period, or calculate a plurality of micro similarity scores based on a plurality of subsets of the full set of the time-series data from the first sensor during the first period and a plurality of subsets of the full set of the time-series data from each sensor in the plurality of associated sensors during the first period, and identify the set of correlated sensors based on the similarity score between the sensor data from the first sensor and the sensor data from the set of correlated sensors exceeding a threshold similarity score, whereby the apparatus is configured to identify the set of correlated sensors, as claimed in claim 15. **Claim 17** the fault in the first sensor includes a predicted fault of the first sensor, and the at least one processor is configured to, based on a sequence prediction model based on a deep learning recurrent neural network, predict a plurality of anomaly scores including at least two of a univariate anomaly score, a bivariate anomaly score, or a multivariate anomaly score, and calculate a predicted fault score indicative of a likelihood of the sensor fault by generating an ensemble of the anomaly scores based on the plurality of anomaly scores, whereby the apparatus is configured to detect the predicted fault of the first sensor, as claimed in claim 16. **Claim 18** A computer-readable medium storing computer-executable code for an apparatus, which when executed by a processor causes the processor to, receive sensor data from a plurality of associated sensors, identify a set of correlated sensors in the plurality of associated sensors for a first sensor in the plurality of associated sensors, and detect a fault in the first sensor based on at least one of the sensor data received from the sensors, the sensor data received from the set of correlated sensors, and the sensor data received from other sensors. A computer-readable medium for implementing a repair strategy based on the detected failure of the first sensor. **Claim 19** When the code is executed by the processor, the processor is caused to detect a failure in the sensor using one or more models for detecting a failure in the sensor based on at least one of the sensor data received from the sensor, the sensor data received from the set of correlated sensors, and the sensor data received from other sensors, wherein the one or more models include a univariate anomaly detection model based on the sensor data received from the first sensor, a bivariate anomaly detection model based on the sensor data received from the first sensor and the sensor data received from the set of correlated sensors, a multivariate anomaly detection model based on the sensor data received from the plurality of related sensors, including the sensor data received from the first sensor, the sensor data received from the set of correlated sensors, and the sensor data received from other sensors, an ensemble anomaly detection model based on the univariate anomaly detection model and the bivariate anomaly detection model, or an ensemble anomaly detection model based on the univariate anomaly detection model and the multivariate anomaly detection model The computer-readable medium according to claim 18, comprising. **Claim 20** The failure in the first sensor includes a predicted failure of the first sensor, and when the code is executed by the processor, the processor is caused to predict a plurality of anomaly scores including at least two of a univariate anomaly score, a bivariate anomaly score, or a multivariate anomaly score based on a sequence prediction model based on a deep learning recurrent neural network, calculate a predicted failure score indicating the likelihood of the sensor failure by generating an ensemble of anomaly scores based on the plurality of anomaly scores, thereby detecting the predicted failure of the first sensor. The computer-readable medium according to claim 18.

Citation Information

Patent Citations

  • Anomaly detection system and method

    JP2016201088A

  • Systems and methods for automated data science processes

    JP2023537766A

  • Sensor fault detection and diagnosis for autonomous systems

    US20160217627A1

  • Systems and methods for an automated data science process

    WO2022039748A1