Real-time detection, prediction, and repair of sensor failures through a data-driven approach.

A data-driven method using anomaly detection techniques addresses sensor failure detection and prediction, ensuring continuous system operation by identifying critical sensors and implementing remediation strategies, reducing costs and human error.

JP7829734B2Active Publication Date: 2026-03-13HITACHI VANTARA LLC
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-04-29
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing sensor failure detection technologies fail to timely detect and predict failures, leading to erroneous readings, system damage, and increased costs due to manual inspections, and often cannot distinguish between sensor failures and operational anomalies.

Method used

A data-driven approach using univariate, bivariate, and multivariate anomaly detection methods, combined with ensemble algorithms, to identify critical sensors, predict failures, and implement remediation strategies based on correlated sensor data, enhancing fault tolerance and reducing manual intervention.

Benefits of technology

Enables real-time detection and prediction of sensor failures, reducing human error and maintenance costs, ensuring continuous system operation by utilizing correlated sensors as substitutes until repairs are made, and providing automated root cause analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007829734000001
    Figure 0007829734000001
  • Figure 0007829734000002
    Figure 0007829734000002
  • Figure 0007829734000003
    Figure 0007829734000003
Patent Text Reader

Abstract

A method for real-time detection, prediction, and repair of sensor failures may include receiving sensor data from a plurality of related sensors. The method may include identifying a set of correlated sensors in the plurality of related sensors for a first sensor in the plurality of related sensors. The method may further include detecting a failure in the first sensor based on at least one of the sensor data received from the first sensor, the sensor data received from the set of correlated sensors, and the sensor data received from other sensors. The method may further include implementing a repair strategy based on the predicted failure of the sensor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to the Internet of Things (IoT) and Operational Technology (OT) domains.

Background Art

[0002] IoT and OT offer great potential to change system functions and business operations by efficiently monitoring and automating systems without the need for human interaction or involvement. IoT and OT systems rely on large amounts of data collected by one or more sensors to automate system operation and decision-making in some applications. Sensors may be devices that respond to inputs from the physical world, capture the inputs, and transmit them to a storage device in some embodiments.

[0003] As used herein, a sensor is a device designed to respond to and / or monitor a particular type of condition in the physical world and then generate a signal (usually an electrical signal) that can represent the magnitude of the monitored condition. As the use of IoT devices and OT applications expands, different types of sensors are used, resulting in different types of data for analysis and processing. In some embodiments, the sensor may include any one or more of a temperature sensor, a pressure sensor, a vibration sensor, an acoustic sensor, a motion sensor, a level sensor, an image sensor, a proximity sensor, a water quality sensor, a chemical sensor, a gas sensor, a smoke sensor, an infrared (IR) sensor, an acceleration sensor, a gyro sensor, a humidity sensor, an optical sensor, and / or a Light Detection and Ranging (LIDAR) sensor.

[0004] Sensor data collected from different types of sensors may be represented differently. For example, some sensors may be analog sensors that capture continuous values ​​and attempt to identify any nuances of what is being measured, or digital sensors that may use sampling to encode what is being measured. As a result, the captured data may be either "analog data" or "digital data." Thus, the data may be numerical, image, or video. In addition, some sensors collect data in a streaming manner and use time-series data to represent the collected values. Other sensors may collect data at isolated points in time.

[0005] IoT and OT industrial systems, in some aspects, rely on sensors that function to monitor the system and collect accurate data for processing, analysis, and modeling in a set of downstream applications. Data quality from sensors plays a fundamental role in the IoT and OT domain in some aspects. Due to the nature of deployment (which can be outdoors and / or in harsh environments) and the limitations of low-cost components, sensors can be prone to failure. In some aspects, a significant portion of failures can result from misalignment and sudden failures in the sensor's sensing component, leading to serious data inaccuracies. As a result, IoT sensors can become misaligned, non-functional, and unreliable, and may output misleading data after operating for some time. In IoT / OT systems, sensors may be installed on assets and connected to storage and / or computing servers via a network for data collection and processing. Any part of the hardware or software used to support the operation of the sensor may fail, resulting in erroneous sensor readings. Failures can occur in the root layer (sensors), network layer (network connectivity), computing layer, or storage layer. While detecting faults at all layers is useful for the accurate and continuous operation of IoT / OT systems, this disclosure focuses on faults in sensors, including direct links to sensors (part of the network layer).

[0006] Currently, scheduled inspections cannot capture faulty sensors in a timely manner, and unnecessary inspections incur additional costs. Such manual inspections are prone to errors and can be time-consuming. This disclosure addresses automated, data-driven approaches for detecting and predicting sensor failures. In addition, root cause analysis is performed for individual failures, and a systematic fault tolerance strategy can be designed to enable the IoT system to continue operating without interruption despite the failure of one or more sensors. In some embodiments, based on detected faulty sensors, the system may identify and perform remedial actions to repair or replace the sensor to avoid any erroneous decisions based on readings from such faulty sensors. [Overview of the Initiative] [Means for solving the problem]

[0007] The exemplary implementations described herein include innovative methods. The method may include receiving sensor data from a plurality of associated sensors. The method may further include identifying a set of correlated sensors in the plurality of associated sensors for a first sensor in the plurality of associated sensors. The method may include detecting a fault in the first sensor based on at least one of the sensor data received from the sensor, the sensor data received from the set of correlated sensors, and the sensor data received from other sensors. The method may further include implementing a remediation strategy based on the detected fault in the first sensor.

[0008] The exemplary implementations described herein include an innovative computer-readable medium for storing computer-executable code. The computer-executable code may include instructions for receiving sensor data from a plurality of associated sensors. The computer-executable code may also include instructions for identifying a set of correlated sensors in a plurality of associated sensors for a first sensor in the plurality of associated sensors. The computer-executable code may further include detecting a fault in the first sensor based on at least one of the sensor data received from the sensor, the sensor data received from the set of correlated sensors, and the sensor data received from other sensors. The computer-executable code may also include instructions for implementing a remediation strategy based on the detected fault in the first sensor.

[0009] The exemplary implementations described herein include an innovative device. The device may include memory and at least one processor configured to collect a set of physical sensor data. The at least one processor may also be configured to receive sensor data from a plurality of associated sensors. The at least one processor may be further configured to identify a set of correlated sensors in the plurality of associated sensors for a first sensor in the plurality of associated sensors. The at least one processor may also be configured to detect a fault in the first sensor based on sensor data received from the sensor, sensor data received from the set of correlated sensors, and sensor data received from other sensors. The at least one processor may be configured to implement a remediation strategy based on the detected fault in the first sensor.

[0010] The exemplary implementations described herein include innovative devices. The devices may include means for receiving sensor data from a plurality of associated sensors. The devices may further include means for identifying a set of correlated sensors in a plurality of associated sensors for a first sensor in the plurality of associated sensors. The devices may also include means for detecting a fault in the first sensor based on at least one of the sensor data received from the sensor, the sensor data received from the set of correlated sensors, and the sensor data received from other sensors. The devices may further include means for implementing a remediation strategy based on the detected fault in the first sensor. [Brief explanation of the drawing]

[0011] [Figure 1] Figure 1 shows a solution architecture for fault detection, fault prediction, fault repair, and fault tolerance.

[0012] [Figure 2A] Figure 2A shows the steps used to identify sensors that enable fault tolerance.

[0013] [Figure 2B] Figure 2B is a diagram for bootstrapping macro-similarity scores.

[0014] [Figure 3A] Figure 3A shows the calculation of the microsimilarity score.

[0015] [Figure 3B] Figure 3B shows a method for bootstrapping the microsimilarity score.

[0016] [Figure 4A]FIG. 4A shows a method of bivariate analysis for identifying whether related and / or corresponding sensors have experienced (or are experiencing) similar problems indicating malfunction, or have not experienced or are not experiencing similar problems indicating that the sensors are malfunctioning.

[0017] [Figure 4B] FIG. 4B is a diagram showing a first ensemble approach for sensor fault detection.

[0018] [Figure 5] FIG. 5 is a diagram showing a second ensemble approach for sensor fault detection.

[0019] [Figure 6] FIG. 6 is a diagram showing an example of a fault prediction module.

[0020] [Figure 7] FIG. 7 is a flowchart of a method for detecting and repairing faults in sensors associated with a system.

[0021] [Figure 8] FIG. 8 is a further enlarged diagram of sub-operations performed to identify a set of correlated sensors in multiple related sensors in some embodiments.

[0022] [Figure 9] FIG. 9 shows an example of an operating environment having an example of a computer device suitable for use in some implementation examples.

BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The following detailed description provides further details of the drawings and implementation examples of this application. For clarity, redundant reference numbers and descriptions of elements between the drawings have been omitted. Terms used throughout this description are provided as examples and are not intended to be limiting. For example, the use of the term “automatic” may include a fully automatic implementation or a semi-automatic implementation that includes user or administrator control over certain aspects of the implementation, depending on the desired implementation for those skilled in the art practicing the implementation of this application. Selection may be made by the user via a user interface or other input means, or may be implemented by a desired algorithm. The implementation examples described herein may be used individually or in combination, and the functionality of the implementation examples may be implemented by any means depending on the desired implementation.

[0024] This disclosure presents systems, apparatus, and methods that address problems associated with conventional sensor failure detection technologies. For example, conventional approaches may fail to detect sensor failures in a timely manner, which can lead to erroneous readings, damage to the system, inaccurate insights, and erroneous decisions. Furthermore, conventional approaches may detect failures after they have already occurred, and therefore may not be able to repair or prevent them. While some approaches use conventional time-series forecasting techniques to predict anomalies in the data, existing approaches may fail to distinguish sensor failures from operational anomalies in the underlying system. Manual inspection of IoT sensors can be error-prone and time-consuming. For example, scheduled inspection of IoT sensors may not capture faulty sensors in a timely manner, thus increasing the risk of obtaining incorrect sensor readings. Unnecessarily aggressive inspection schedules designed to mitigate the risk of incorrect sensor readings may be associated with additional unnecessary costs.

[0025] This disclosure presents systems, apparatus, and methods that provide techniques relating to detecting failures in sensors associated with a system (e.g., IoT sensors associated with industrial and / or manufacturing systems and / or processes). The method may include receiving sensor data from a plurality of associated sensors. The method may further include identifying a set of correlated sensors in the plurality of associated sensors for a first sensor in the plurality of associated sensors. The method may also include detecting a failure in the first sensor based on sensor data received from the sensor, sensor data received from the set of correlated sensors, and sensor data received from other sensors. The method may further include implementing a remediation strategy based on the detected failure in the first sensor.

[0026] In general, to address some of the problems identified above, the method may involve one or more of the following: critical sensor identification, fault tolerance identification, fault detection, fault prediction, fault repair, and / or fault tolerance. For example, critical sensor identification may include identifying one or more critical sensors based on domain knowledge, data analysis, or downstream tasks. Critical sensors may include sensors that capture critical data to monitor the health of the underlying system, and / or sensors that capture critical data used in at least one of the following: identification of repair strategies, derivation of business insights, or construction of solutions for a set of related downstream tasks such as anomaly detection, fault prediction, and remaining service life prediction.

[0027] Fault tolerance identification may, in some embodiments, include identifying a set of one or more correlated sensors for each critical sensor (e.g., a target sensor). The identified correlated sensors may include one or more sensors that capture similar or highly correlated signals based on similarity metrics, and some approaches may be used to compute one or more similarity scores (e.g., similarity scores associated with similarity metrics) between data from two sensors. Fault detection may, in some embodiments, include detecting faults in one or more sensors based on data from physical sensors and / or data associated with virtual sensors (e.g., expected data for virtual sensors based on relevant data from one or more physical sensors processed using one or more physical-based models). Fault detection may, in some embodiments, include one or more univariate anomaly detection, bivariate anomaly detection, and / or multivariate anomaly detection approaches, and may further involve an ensemble algorithm based on one or more univariate, bivariate, or multivariate anomaly detection approaches.

[0028] In some embodiments, failure prediction may include predicting failure in one or more sensors (e.g., critical sensors). Failure prediction may be based on a deep learning recurrent neural network (RNN) model using sensor data from at least the critical sensor and additional sensors (e.g., a set of related and / or correlated sensors). Based on failure prediction, in some embodiments, the method may include identifying fault repair and / or fault tolerance actions. Fault repair actions (e.g., repairing or replacing the sensor predicted to fail) may be based on root cause analysis related to failure prediction, and further on domain knowledge indicating one or more remedial actions based on the results of the root cause analysis. In some embodiments, failure prediction for a particular sensor under consideration allows the system (or method) to identify (or obtain the identity of) a set of correlated sensors that can be used in place of the particular sensor under consideration until the sensor is repaired or replaced.

[0029] As will be described in more detail below, the apparatus and methods described herein may provide a data-driven approach for sensor failure detection that can distinguish between sensor failures and operational failures in the underlying system, based on a novel combination of univariate anomaly detection and bivariate / multivariate anomaly detection. In some embodiments, the data-driven approach uses data from multiple sensors associated with the system (e.g., a set of correlated and / or related sensors) to detect and / or predict failures of a particular sensor of interest. A data-driven approach that includes one or more critical sensor identification, fault tolerance identification, fault detection, fault prediction, fault repair, and / or fault tolerance may, in some embodiments, enable the identification of fault tolerance actions / opportunities, such as (1) repairing faults before they occur based on fault sensors to avoid damage to unmonitored underlying systems by providing fault prediction and fault detection; (2) reducing manual costs and / or undetected faults associated with maintenance schedules by providing real-time fault detection; (3) reducing human error in sensor maintenance and diagnostics by relying on data; (4) drawing conclusions based on existing evidence; and (5) relying on data collected by correlated sensors (or “digital twin models” of the sensors in question) until the sensor in question is repaired or replaced, based on the fault tolerance identification operations provided herein. In some embodiments, the data-driven approach may rely on both current and historical data from a set of sensors in question and correlated sensors.

[0030] The system, apparatus, and / or method may provide automated root cause analysis. Root cause analysis of failures may be performed manually based on domain knowledge and data visualizations, which can be subjective, time-consuming, and error-prone. In some cases, the root cause may be associated with raw sensor data not addressed by the domain knowledge or data visualizations used in manual root cause analysis. The system, apparatus, and / or method may provide automated root cause analysis based on a standardized approach to identifying the root cause of a predicted failure.

[0031] In some embodiments, sensor data (e.g., IoT sensor data, vibration data) may be high-frequency data (e.g., 1000Hz to 3000Hz). High-frequency data, in some embodiments, presents challenges to constructing solutions for failure prediction problems. For example, high-frequency data may be associated with high levels of noise or long or resource-intensive analysis (e.g., computation) times. The sampling frequency or aggregation window may require optimization to accurately predict either short-term or long-term failures. Accordingly, the system, apparatus, and / or method may provide a window optimization operation that identifies optimized windows and / or aggregate statistics for failure prediction.

[0032] In some embodiments, physical sensor data may not capture all signals that could be useful for monitoring a system due to the harsh environment in which the sensors are installed, the cost of the sensors, and / or the functionality of the sensors. As a result, the collected data may not be sufficient to monitor the health of the system and capture potential risks and failures. In some embodiments, the inability to capture all potentially useful signals may pose a challenge to the construction of failure prediction solutions. Therefore, the system, apparatus, and / or method may enhance physical sensor data to capture the signals necessary to assist in the construction of system monitoring and failure prediction solutions. For example, physical sensor data may be processed by a set of physical-based models to generate virtual sensor data.

[0033] Figure 100 shows a solution architecture for fault detection, fault prediction, fault repair, and fault tolerance. The solution architecture may include a sensor data module 110. The sensor data module 110 may incorporate a set of physical sensors 110a and a set of virtual sensors 110b. The physical sensors 110a may include any one or more of the following: temperature sensors, pressure sensors, vibration sensors, acoustic sensors, motion sensors, level sensors, image sensors, proximity sensors, water quality sensors, chemical sensors, gas sensors, smoke sensors, IR sensors, acceleration sensors, gyroscopes, humidity sensors, light sensors, and / or LIDAR sensors. Data from the physical sensors 110a may be provided to a set of physical-based models to generate data associated with the virtual sensors 110b.

[0034] In some embodiments, the physical sensor 110a may be installed on a target asset (e.g., an asset in an OT system) and may be used to collect data for monitoring the health and performance of the asset. Different types of sensors are designed to collect different types of data across different industries, different assets, and different tasks. Different sensors may be included in the physical sensor 110a for different applications, but this disclosure will discuss them generally because the method can be applied to a wide range of sensors and data types. In some embodiments, the virtual sensor 110b may be associated with output variables from a set of physical base models or a set of digital twin models, which can complement and / or verify data from the physical sensor and thus help monitor and maintain system health. In the case of “complementary,” if physical sensor data is not available or insufficient, the virtual sensor data from the digital twin model may function as a “substitute” for the physical sensor. In the case of "verification," assuming that physical sensors also collect data as output of the digital twin model, the virtual sensor data can function as "expected" values, while the values ​​from the physical sensors can function as "observed" values. Therefore, the variance or difference between them can be used as a signal to detect abnormal behavior in the system.

[0035] Data collected by one sensor S1 may, in some embodiments, be closely related to data collected by another sensor S2. In this case, S1 can be a substitute for S2, and vice versa. For example, wind turbine shaft torque can be approximately represented by the amount of vibration generated by the generator, and vice versa. Such substitution relationships can be obtained based on domain knowledge and / or data analysis (correlation analysis, etc.). Substitute sensors enable fault tolerance, and if one sensor fails, the other sensor may be used as a substitute to construct a solution. In addition, a faulty sensor may be recognized when such a substitution relationship does not hold.

[0036] Sensor data from the sensor data module 110 may be provided to the critical sensor identification module 120. In some embodiments, multiple sensors may be installed on assets within an industrial system to monitor the health of the system, but only some of the sensors are useful for deriving insights, making decisions, and / or constructing solutions for downstream tasks. Such sensors are essential for maintaining a healthy, reliable, and continuously operating industrial system, and it is necessary to keep these sensors functioning as expected. Such sensors may be referred to as critical sensors, and special attention may be paid to these critical sensors beyond that paid to other non-critical sensors in the system. There are several approaches to identifying critical sensors that can be employed by the critical sensor identification module 120.

[0037] In some embodiments, a domain knowledge-based approach may be used. For example, operators and / or technicians may possess domain knowledge that often enables them to identify which sensors are useful and essential for monitoring the health of the system. In some embodiments, operators and / or technicians may provide input to a critical sensor identification module 120 to identify critical sensors. For example, that domain knowledge may be used to identify a list of critical sensors ranked by their importance.

[0038] A data-driven approach may be used to identify one or more critical sensors. For example, one or more variables used to indicate the health of a system may be identified, and data analysis may be performed on historical data from multiple sensors associated with the system to identify which sensors are closely related to the health indicator variables. One approach is to calculate the correlation coefficient between the data from each sensor and the health indicator variables. As a result, a list of sensors ranked by that correlation coefficient can be obtained.

[0039] Sensor data may be used to construct solutions for several downstream tasks, including but not limited to failure prediction, anomaly detection, remaining service life, and yield optimization. Since solutions for downstream tasks are constructed based on data from multiple sensors, the importance of sensors for such solutions can be identified through one or more downstream task-based approaches (e.g., feature selection techniques and / or explainable AI techniques). Based on models constructed on downstream tasks, the system and / or method may calculate values ​​that reflect the feature importance of each sensor (e.g., values ​​relating to the explanatory impact of sensor data on downstream tasks). One or more downstream task-based approaches may, in some embodiments, provide a list of sensors ranked by their importance.

[0040] The approaches described above, such as a domain knowledge-based approach, a data-driven approach, and a downstream task-based approach, may be used independently to identify critical sensors. In some embodiments, different approaches may be combined into a single approach by merging ordered lists of sensors generated by the different approaches. For example, one possible approach is to calculate the average ranking of each sensor based on its ranking in three lists, and then sort the sensors based on their average ranking. When calculating the average ranking, a weighted average can be used by first assigning weights to each approach, and then the average rank can be calculated using the weighted rank.

[0041] The fault tolerance identification module 130 may be used to identify one or more sensors that can provide alternative sensor data for a critical sensor. For example, given sensor S1, a set of one or more sensors that can function as a substitute for sensor S1 may be identified. Based on the set of one or more sensors identified by the fault tolerance identification module 130 and the predicted or detected failure of sensor S1, the set of one or more "substitute" sensors may be used in place of sensor S1 for at least a period of time until S1 is repaired or replaced.

[0042] Figure 2A is a diagram showing the steps used to identify sensors that enable fault tolerance. In 210, the method may acquire data from all sensors and take the time-series sensor data values ​​as vectors. The sensors may be physical sensors and / or virtual sensors from a digital twin model in some embodiments. In 220, the method may calculate a similarity score between two vectors for each pair of sensors. To make a comparison, in some embodiments, the data from the pairs of sensors is normalized so that they can be compared. For example, the data from the pairs of sensors may have been collected initially over different time periods or at different frequencies, and the method may sample data from different sensors to make the time window and data frequency the same. Once the data is normalized, the method may select one or more similarity metrics. One or more similarity metrics may include, but are not limited to, the correlation coefficient, cosine similarity, Hamming distance, Euclidean distance, Manhattan distance, and / or Minkowski distance. Based on one or more selected similarity metrics, the method may measure the similarity between two vectors. Calculating a similarity score between two vectors for a pair of sensors may involve one of several approaches. Theoretically, the method may calculate a similarity score for each possible pair of sensors (e.g., physical and / or virtual sensors), but in some embodiments, the similarity score may be calculated for each of the critical sensors and other sensors (including both critical and non-critical sensors).

[0043] The similarity score calculated in 220 may be compared in 230 to a threshold similarity score to determine whether two sensors are correlated and / or related. If the calculated similarity score is above the threshold, the similarity may be verified based on domain knowledge by the operator and / or technician. The similarity identified in 230 may be referred to as macro similarity, based on a comparison of larger datasets (e.g., data collected over a day, a week, etc.) than those used to identify micro similarity, as described below with respect to Figures 3A and 3B. As described above, macro similarity may be used to identify a set of "substitute" sensors for a critical sensor or a set of related and / or correlated sensors for repair or other downstream tasks.

[0044] Figure 2B is Figure 205 for bootstrapping the macrosimilarity score. If the two data vectors from the two sensors exceed a threshold length, the similarity calculation can require significant resources and take an extremely long time to complete. Figure 205 shows a workflow for how to use the bootstrapping technique to calculate the similarity score. Figure 205 shows that the method may include acquiring data at 240 as described above with respect to step 210 in Figure 200.

[0045] After acquiring the data, the method may specify whether to analyze the complete dataset or a reduced (e.g., bootstrapped) dataset (not shown). The reduction (e.g., bootstrapping) technique may include sampling corresponding data from each of two vectors at 250. For example, the method may sample from two vectors at 250 by substitution with a predetermined sampling rate, e.g., 0.01, and use this to calculate and / or compute a similarity score at 260. Calculating the similarity score at 260 is similar to calculating the similarity described with respect to 220 in Figure 200, but is performed only in a shorter time window than usual.

[0046] After calculating the similarity score for the current sample, the method may proceed in step 270 to determine whether the threshold number of iterations has been met (each iteration is associated with a sampling-based similarity score). The threshold number of iterations may be set before the analysis and may be chosen to be large enough to ensure that the calculated value is the reported value. If the threshold number of iterations has not been met, the method may return to step 250. Thus, the method may repeat this process multiple times to obtain multiple similarity scores. The method may then aggregate the similarity scores from the multiple runs in step 280 and use the aggregated value as the final similarity score. The aggregation function used in step 280 may include, but is not limited to, the mean, weighted mean, maximum, minimum, median, weighted median, etc. Then, based on the aggregated similarity score, also in step 280, the aggregated similarity score can be compared to a predetermined similarity score threshold to determine whether the two vectors are similar.

[0047] Figure 3A shows the calculation of microsimilarity scores. In some embodiments, instead of calculating a single similarity score (or aggregated similarity score), as described with respect to steps 220 and 280, the method may calculate a series of similarity scores based on data within a time window (or time segment). Figure 3A shows a workflow 300 on how microsimilarity calculation works. With respect to generating macrosimilarity scores in Figure 2A, the method may first acquire data for pairs of sensors between the same (or overlapping) time windows in 310. In 220, the method may identify a strategy used to define the time windows used when calculating microsimilarity scores. In some embodiments, a time window is one of a rolling window or an adjacent window. A time window may also depend on an event. For example, holiday seasons, daily business hours, weekdays, weekends, etc., may be used to identify a time window.

[0048] For each time window, the method may calculate the similarity score in 330, resulting in a set of similarity scores for each pair of sensors. The method may then, in 340, obtain the distribution of similarity scores based on their values ​​and frequencies, and analyze the distribution of similarity scores. To determine whether two sensors are similar, the method may perform a statistical significance test to determine whether a given similarity score threshold is significantly different from the distribution of similarity scores. For example, the method may use a one-sample one-tailed t-test (or other appropriate statistical analysis) to determine whether the similarity score threshold is significantly lower than the similarity score. The method may first calculate a statistic based on the data of the similarity score threshold relative to the distribution of similarity scores. Then, based on the significance level, the method may determine whether the similarity score threshold is significantly lower than the similarity score. In this case, the focus is on a one-tailed test, i.e., the left-side critical region in the distribution of similarity scores. Micro-similarity provides a granular view of similarity scores and therefore provides more information and accuracy to represent the similarity of two sensors.

[0049] Figure 3B is Figure 305, which illustrates a bootstrapping method based on microsimilarity. Similar to the relationship between Figures 2A and 2B, Figure 3B shows that the first two steps of the method, namely acquiring data in 350 and identifying a strategy to define time windows in 355, are equivalent to steps 310 and 320, respectively. In the microsimilarity approach, if there are too many time windows, it can be computationally intensive and take too long to run. Therefore, the method shown in Figure 305 may employ a bootstrapping technique to solve such problems. Once the method has identified a window generation strategy and defined all time windows in 355, in 360, the time windows can be sampled by substitution with a predefined sampling rate, e.g., 0.01, using the bootstrapping technique. The method may then apply the microsimilarity approach in 365 to calculate a set of similarity scores and a distribution of similarity scores. In 370, the method may compare a similarity score threshold to the distribution of similarity scores based on a statistical significance test, and the results of the current run are recorded. In 375, the method determines whether additional runs should be performed. If so, the method may return to 360 and perform another random sample within the time window defined in 355. The sampling runs may continue with several runs of bootstrapping sampling and application of the microsimilarity approach until a predetermined number of runs are reached. The results from the predetermined number of runs may be aggregated in 380 to obtain a final result. Since the results from each run are binary values ​​indicating whether the similarity score is significantly lower than the similarity score (meeting the threshold criterion for identifying similarity via the identity score), some embodiments use a "majority vote" technique to determine which binary value dominates the results and use that as the final result. In other embodiments, if the results from each run are represented by numerical scores indicating statistical significance, the mean statistical significance value can be calculated as the final result using a mean or weighted average technique.Finally, in some embodiments, determining whether two sensors are similar may include verifying the similarity based on domain knowledge by the operator and / or technician if the calculated similarity score exceeds a threshold. Generally speaking, approaches to calculating bootstrapping similarity, micro-similarity, and bootstrapping micro-similarity each transform the original calculation for a large vector into multiple calculations for smaller vectors, which reduces hardware requirements. As a result, the analysis may be able to be performed on edge devices (e.g., devices that may have limited hardware resources) using these approaches.

[0050] In some embodiments, given a sensor S1, there may be no single sensor that can function as a substitute for sensor S1, and the method may select a group of sensors as a whole that can be used as a substitute for sensor S1. One approach is to use sensor S1 as a target and the remaining sensors as features to build a machine learning model. If the model performance metrics exceed a certain threshold, then sensor S1 can be said to be substituted by a set of one or more correlated (or related) sensors. To identify substitute sensors, the method may select important features from the model and use the corresponding sensors as substitute sensors for sensor S1. Feature selection may be based on several feature selection techniques, including but not limited to forward selection, backward selection, and model-based feature selection. Domain knowledge may be incorporated in some embodiments to improve feature selection. The group of sensors used as a substitute for the sensor in question (in this case, S1) is called a cohort sensor, related sensor, or correlated sensor. In addition to using a group of cohort sensors as a substitute for the sensor in question, the output from a machine learning model (or a set of physically based models associated with a virtual sensor) can also be used as a substitute for the sensor in question.

[0051] In addition to identifying similarities between sensors based on sensor data, the methods described herein may incorporate some domain knowledge, where available. For example, some exemplary sensors that may have a high degree of similarity are a) physical sensors for input variables and input variables in the designed motion profile, b) physical sensors for output variables and output variables from a digital twin model (which may be based on input variables from either the designed motion profile or the physical sensors for the input variables), and / or c) output variables from different versions of the digital twin model (which may be based on input variables from either the designed motion profile or the physical sensors for the input variables).

[0052] The methods described with respect to Figures 2A-3B relate to and / or are performed by a critical sensor identification module 120 and / or a fault tolerance identification module 130. Based on the output of the fault tolerance identification module 130 and, in some embodiments, the critical sensor identification module 120, a fault detection module 140 may perform fault detection operations. For example, after a critical sensor is identified by the critical sensor identification module 120 and similar sensors are identified (if possible) for the critical sensor by the fault tolerance identification module 130, one or more data-driven approaches are used to detect faults in the critical sensor. The data may be physical sensor data and / or virtual sensor data from a digital twin. In some embodiments, this approach may be accompanied by one or more machine learning models, including a univariate anomaly detection model, a bivariate anomaly detection model, and / or a multivariate anomaly detection model.

[0053] For the sensors under consideration, univariate anomaly analysis may include running an anomaly detection model on the sensor data. Anomalies in the temporal sequence of data may indicate either a faulty sensor or an operational anomaly. Therefore, in some embodiments, a second anomaly analysis is performed to distinguish between faulty sensors and operational anomalies. For example, Figure 4A shows a bivariate analysis method 400 that identifies whether related and / or corresponding sensors have experienced (or are experiencing) similar problems indicating operational anomalies, or have not experienced (or are not experiencing) similar problems indicating a sensor failure. For example, in 410, for the sensors under consideration, the method first identifies similar sensors based on the output from a fault tolerance identification module 130. For cohort sensors, the method may use the output from a machine learning model based on a group of cohort sensors as similar sensors.

[0054] In step 420, the method may then perform a micro-similarity algorithm on historical data from the sensor of interest and similar sensors, as described above with respect to Figures 3A and 3B above, in order to calculate a series of similarity scores. The method may also obtain a distribution f of similarity scores. After performing the micro-similarity algorithm on the historical data, in step 430, the method performs micro-similarity on new data from the sensor of interest and similar sensors to obtain the current similarity scores.

[0055] Finally, in 440, the method may check whether the similarity score based on historical data differs from the similarity score based on current data to a degree that indicates a faulty sensor. For example, an anomaly detection model can be run on a set of similarity scores to detect such differences or anomalies, and / or the method can perform a statistical significance test of the similarity scores against the distribution of similarity scores. A one-sample t-test may be performed by selecting a significance level, e.g., 0.01. Anomalies detected by a bivariate anomaly detection model typically indicate a fault in the sensor in question or a similar sensor.

[0056] In some embodiments, data from multiple sensors may be used to construct a multivariate anomaly detection model. Such anomalies typically indicate system malfunctions, assuming that it is unlikely that multiple sensors will fail simultaneously. Of the three approaches described above, the univariate anomaly detection model may not be able to distinguish between anomalies caused by sensor failures or system malfunctions, the bivariate anomaly detection model may not be able to determine which of two sensors is faulty, and the multivariate anomaly detection model detects only system malfunctions. To address these limitations, some approaches use an ensemble or combination of the above approaches to determine which sensor is faulty. We introduce two approaches for this purpose. Each approach can be performed independently to detect sensor failures.

[0057] Figure 4B is a diagram of Figure 405 illustrating a first ensemble approach for sensor failure detection. In the first ensemble approach shown in Figure 405, the outputs from a univariate anomaly detection model performed at 450 and a bivariate anomaly detection model performed at 470 may be used to detect sensor failures. This approach leverages the existence of related or corresponding sensors for the sensor of interest. As described above, similar (e.g., related or correlated) sensors may be detected by the fault tolerance identification module 130. For example, referring to Figure 4B, the first ensemble approach may include running a univariate anomaly detection model at 450 on a vector of sensor data for a specific sensor of interest (e.g., a critical sensor). Based on the univariate anomaly detection model performed at 450, the method may determine at 460 whether an anomaly was detected. If no anomaly was detected, the sensor may be determined at 460 to be not faulty at 490B. However, if an anomaly is detected by the univariate anomaly detection model performed at 450, the method may perform a bivariate sensor anomaly detection model at 470 on the vectors of the sensor in question and the associated / correlated / similar sensors. Based on the bivariate anomaly detection model performed at 470, the method may determine at 480 whether an anomaly (e.g., an anomaly between the outputs of the sensor in question and the associated / correlated / similar sensors) has been detected. If no anomaly is detected (e.g., the associated / correlated / similar sensors produce similar measurements / vectors to those of the sensor in question that are produced by the sensor), the sensor in question may be identified at 480 as a non-faulty 490B. However, if an anomaly is detected based on the bivariate anomaly detection model performed at 470, the sensor in question may be identified at 480 as a faulty 490A, because there is evidence that the data from the sensor in question (or critical sensor) does not match the sensor data collected by the associated / correlated / similar sensors, and the univariate anomaly detection model can identify an anomaly in the sensor in question and conclude that the sensor in question is faulty.

[0058] Figure 500 shows a second ensemble approach for sensor failure detection. In the second ensemble approach shown in Figure 500, the outputs from a univariate anomaly detection model performed at 550 and a multivariate anomaly detection model performed at 570 may be used to detect sensor failures. The second ensemble approach may include performing a univariate anomaly detection model at 550 on a vector of sensor data for a specific sensor of interest (e.g., a critical sensor). Based on the univariate anomaly detection model performed at 550, the method may determine at 560 whether an anomaly was detected. If no anomaly was detected, the sensor may be determined at 560 to be not faulty at 590B. However, if an anomaly was detected by the univariate anomaly detection model performed at 550, the method may perform a multivariate sensor anomaly detection model on a vector for all (critical) sensors (including the sensor of interest). Based on the multivariate anomaly detection model performed at 570, the method may determine at 580 whether an anomaly was detected. If an anomaly is detected, the sensor may be identified in 580 as not having failed in 590B (for example, if a multivariate anomaly detection model based on all (critical) sensors detects an anomaly, it is likely a system / operational anomaly and not a sensor failure). However, if no anomaly is detected based on the multivariate anomaly detection model run in 570, the multivariate anomaly detection model does not identify any system / operational anomaly and is therefore likely a sensor failure, so the sensor may be identified in 580 as having failed in 590A.

[0059] In some embodiments, additional considerations may be used to identify a faulty sensor. For example, a sensor may be identified as faulty if it is unable to produce any data reading. In some embodiments, the first and second ensemble approaches may be performed simultaneously to detect a faulty sensor. If both approaches detect a fault in the sensor, the sensor may be identified as faulty. If neither approach detects a fault in the sensor, the sensor may be identified as not faulty. If one approach detects a fault in the sensor and the other does not, either a faulty or non-faulty sensor is output, depending on the risk-cost trade-off.

[0060] When constructing an anomaly detection model, time series data may, in some cases, be preprocessed before applying the model to the preprocessed data. Some preprocessing techniques may include, but are not limited to, differences, moving averages, moving variances, window-based features, etc. The approach can be applied to both analog and digital sensors. For digital sensors, the data can first be preprocessed by moving averages and / or moving variances, and then the data becomes continuous.

[0061] Anomaly detection may, in some embodiments, be a distribution-based method. For example, a moving variance based on sensor data may be calculated first, and then the distribution of the moving variance may be calculated or identified. Based on the distribution of the moving variance, anomaly detection (e.g., performed by an anomaly detection module) may identify outliers / anomalies based on a predetermined threshold (e.g., outside the 99% range). Such outliers / anomalies may, in some embodiments, be identified as corresponding to a faulty sensor. The assumption here is that if the sensor data remains at the same value for a period of time (i.e., the moving variance is close to 0), then the deviation from that value corresponds to a sensor failure.

[0062] Fault detection techniques may detect faults when they occur. While a faulty sensor is being repaired and / or replaced, the underlying system may be left unmonitored due to sensor downtime. To avoid the underlying system being left unmonitored during maintenance / repair, in some embodiments, a fault prediction module is provided to predict sensor failures in advance to avoid them, or to enable repair without system unmonitored downtime. Figure 6 is Figure 600 showing an exemplary fault prediction module 650. In some embodiments, the fault prediction module 650 corresponds to the fault prediction module 150 in Figure 1. The fault prediction module 650 may run a set of anomaly detection models 651a and / or 652a to obtain a corresponding set of anomaly scores. The set of anomaly detection models may include one or more univariate anomaly detection models for each sensor's data, a bivariate anomaly detection model for each pair of identified related / correlated / similar sensors, or a multivariate anomaly detection model to generate a set of anomaly scores, as described above.

[0063] The failure prediction module 650 may identify and / or prepare features associated with a set of sensors (e.g., associated with sensor data 651b) (e.g., in the feature module 651). For example, for each sensor, the failure prediction module 650 may obtain the following data: (1) data from the sensor and similar sensors (if available), (2) an anomaly score from a univariate anomaly detection model, (3) an anomaly score from a bivariate anomaly detection model, and (4) an anomaly score from a multivariate anomaly detection model. The failure prediction module 650 may also prepare a target for each sensor (e.g., in the target module 652), and preparing a target may include obtaining one or more of the following for the associated lead time: (1) an anomaly score from a univariate anomaly detection model, (2) an anomaly score from a bivariate anomaly detection model, and / or (3) an anomaly score from a multivariate anomaly detection model, where the lead time may be a predetermined value indicating how far in advance an anomaly is predicted. Based on the acquired data, the failure prediction module 650 may construct a sequence prediction model (e.g., failure prediction model 653). In some embodiments, the sequence prediction model may be a deep learning recurrent neural network (RNN) used for sequence prediction. The deep learning RNN may be either a long short-term memory (LSTM) model or a gated recurrent unit (GRU) model. In some embodiments, the deep learning RNN model may allow multiple targets at once such that the output from each prediction includes three anomaly scores: univariate, bivariate, and multivariate anomaly scores (e.g., associated with the predicted anomaly score 655). Finally, the ensemble approach described above with respect to Figures 4B and 5 may be applied by an ensemble prediction anomaly score module 657 to predict whether a sensor is malfunctioning.

[0064] The fault repair and fault tolerance module 160 may receive information from one or more of the critical sensor identification module 120, the fault tolerance identification module 130, and / or the fault prediction module 150. When a fault is predicted, the fault repair and fault tolerance module 160 may identify the root cause of the fault using explainable AI techniques (such as ELI5 and Shapley Additive Explanation (SHAP)). In this case, the fault repair and fault tolerance module 160 may identify the data point (or critical sensor) and / or root cause associated with the predicted fault. Post-processing of the predicted anomaly score and root cause may verify the fault with or without applying domain knowledge. If the fault is valid, the fault repair and fault tolerance module 160 may identify one or more repairs (including identifying a “substitute” sensor) that can be performed by an operator or technician. For example, identifying one or more repairs may include checking whether the faulty sensor has a set of related / correlated / similar sensors to enable fault tolerance. If a set of related / correlated / similar sensors exists, one or more identified remediations may include using the set of related / correlated / similar sensors for downstream tasks and replacing the faulty sensor. If a set of related / correlated / similar sensors does not exist and / or cannot be identified, one or more identified remediations may include immediately replacing the faulty sensor and adding at least one additional sensor to provide redundancy in the future. If one or more identified remediations exist consecutively (e.g., upstream or downstream of the faulty sensor), geolocation-based faulty sensor remediation may include using upstream and / or downstream sensors to supplement the faulty sensor values. If, for any reason, a sensor generates fault values ​​only over a specific period, time-based faulty sensor remediation may also include using pre- and post-failure data to supplement the sensor values ​​during the failure period.

[0065] In some embodiments, fault tolerance may be systematically implemented during design time and / or operation time. For example, a digital twin model may be constructed to output virtual sensor data as a substitute for fault tolerance to physical sensor data. In some embodiments, virtual sensors from the digital twin model may help complement and validate physical sensors. Two or more versions of the digital twin model may be constructed in some embodiments to enable more fault tolerance and data validation. In some embodiments, the implementation of fault tolerance may include identifying critical sensors and, if similar sensors are not available, introducing at least one additional sensor for fault tolerance. For example, in a service-oriented architecture (SOA), it may be desirable to ensure fault tolerance to critical sensors.

[0066] Figure 7 is a flowchart 700 of a method for detecting and repairing a failure of a sensor associated with a system. The method may be performed by a system such as the solution architecture shown in Figure 100, or by one or more components of individual or combined solution architectures, such as a sensor data module 110, a critical sensor identification module 120, a fault tolerance identification module 130, a fault detection module 140, a fault prediction module 150 (or 650), and / or a fault repair and fault tolerance module 160. In 710, the method may receive sensor data from multiple associated sensors. In some embodiments, the multiple associated sensors are sensors that monitor the same system. In some embodiments, the multiple associated sensors include at least one physical sensor installed in the system or a virtual sensor derived from a set of physical sensors based on a physical-based model. For example, sensor data may be received from multiple associated sensors by the sensor data module 110, or received from the sensor data module 110 by another module of the solution architecture.

[0067] In 720, the method may identify a set of correlated sensors in a plurality of related sensors for a first sensor in a plurality of related sensors. The first sensor is, in some embodiments, a first critical sensor, which captures critical data for monitoring the health of the underlying system and may be used for at least one of identifying a remediation strategy, deriving business insights, or constructing solutions for a set of downstream tasks. The downstream tasks may, in some embodiments, include one or more of anomaly detection, failure prediction, and remaining service life prediction. In some embodiments, the set of correlated sensors includes a set of sensors whose outputs correlate with the output of the first sensor. In some embodiments, the set of correlated sensors includes multiple sensors in a plurality of related sensors. For example, in some embodiments, the correlation may be calculated between critical sensor data and a first principal component of the plurality of sensor data (or several other features of the plurality of sensor data even if the individual sensor data do not correlate at a threshold level). As shown in the enlarged view of 720, in order to identify a set of correlated sensors in a plurality of related sensors in 720, in some embodiments, the method may further calculate a similarity score in 720A between sensor data from a first sensor and sensor data from sensors in a plurality of related sensors, and then identify a set of correlated sensors in 720B based on the similarity score between sensor data from a first sensor and sensor data from a set of correlated sensors exceeding a threshold similarity score.

[0068] Figure 8 is a further extension of the diagram of suboperations performed in 720 to identify a set of correlated sensors in multiple related sensors in several embodiments. Elements 820, 820A, and 820B in Figure 8 correspond to elements 720, 720A, and 720B in Figure 7 in several embodiments, and may further identify suboperations associated with calculating similarity scores in 720A / 820A. In 820A-1, the method may calculate a macro-similarity score based on a full set of time-series sensor data from a first sensor during a first period and a full set of time-series data from each sensor in multiple related sensors during the first period. In addition, in 820A-2, the method may calculate multiple micro-similarity scores based on multiple subsets of the full set of time-series data from the first sensor during a first period and multiple subsets of the full set of time-series data from each sensor in multiple related sensors during the first period. As explained above regarding Figures 2A to 3B, the calculations may be "standard" as in Figures 2A and 3A, or they may be "bootstrapped" as in Figures 2B and 3B.

[0069] In 720B / 820B, the method may identify a set of correlated sensors based on the similarity score between sensor data from a first sensor and sensor data from a set of correlated sensors exceeding a threshold similarity score. The set of correlated sensors may then be identified as a set of one or more "substitute" sensors for a critical sensor in the event of sensor failure. For example, 720 may be performed by a fault tolerance identification module 130.

[0070] In 730, the method may detect a fault in the first sensor based on at least one of sensor data received from the sensor, sensor data received from a set of correlated sensors, and sensor data received from other sensors. Fault detection in 730 may be performed by a fault detection module 140 or a fault prediction module 150. In some embodiments, detecting a sensor fault in 730 based on at least one of sensor data received from the sensor, sensor data received from a set of correlated sensors, and sensor data received from other sensors includes, in 730A, using one or more of the following models to detect a sensor fault: (1) a univariate anomaly detection model based on sensor data received from the first sensor, (2) a bivariate anomaly detection model based on sensor data received from the first sensor and sensor data received from a set of correlated sensors, and / or (3) a multivariate anomaly detection model based on sensor data received from all sensors, including the first sensor, sensor data received from a set of correlated sensors, and other sensors in the system. In some embodiments, the method may also use one or more of the following ensemble models in 730B to detect a fault in the sensor in 730: (1) an ensemble anomaly detection model based on a univariate anomaly detection model and a bivariate anomaly detection model, or (2) an ensemble anomaly detection model based on a univariate anomaly detection model and a multivariate anomaly detection model.

[0071] In embodiments for detecting real-time failures in the first sensor, the univariate (anomaly) score, the bivariate (anomaly) score, and / or the multivariate (anomaly) score may be calculated in 730A based on a univariate anomaly detection model, a bivariate anomaly detection model, and / or a multivariate anomaly detection model, respectively. In some embodiments, the ensemble (anomaly) score may be calculated in 730B based on either (1) an ensemble anomaly detection model based on the univariate and bivariate anomaly scores, and / or (2) an ensemble anomaly detection model based on the univariate and multivariate anomaly scores, or both. In some embodiments, detecting a failure in the first sensor in 730 includes detecting a predicted failure in the first sensor. In some embodiments, detecting a predicted failure in the first sensor may be based on a sequence prediction model based on a deep learning RNN, as discussed above. In some embodiments, the sequence prediction model may be either an LSTM or a GRU model. As part of detecting a predicted failure in the first sensor in 730, the method may, in 730A, calculate (or predict) a plurality of predicted anomaly scores, including at least two of univariate anomaly scores, bivariate anomaly scores, or multivariate anomaly scores, based on a prediction model (e.g., the sequence prediction model described above). In 730B, the method may further calculate a predicted failure score indicating the likelihood of sensor failure by generating an ensemble of anomaly scores based on the plurality of anomaly scores calculated in 730A and one or both of an ensemble anomaly detection (prediction) model based on (1) univariate and bivariate anomaly scores, and / or (2) univariate and multivariate anomaly detection (prediction) models. For example, 730, 730A, and 730B may be performed by a failure prediction module 150.

[0072] Finally, in 740, the method may implement a repair strategy based on the failure of the first sensor detected in 730 (e.g., initiating, suggesting to the user to implement, etc.). For example, 740 may be implemented by a fault repair and fault tolerance module 160. In some embodiments, the repair strategy includes replacing sensor data from the first sensor with sensor data from one or more sensors in a set of correlated sensors. In some embodiments, the repair strategy is based on root cause analysis of the detected fault based on one or more explainable artificial intelligence (AI) techniques (e.g., ELI5 or SHAP). Given prediction results from a machine learning model, the explainable AI technique can discover the contribution of each feature (used in the machine learning model) to the prediction results. The contribution is measured by weight values ​​that can be used by operators and / or technicians to uncover the root cause of the prediction results. Sensor data from one or more sensors within a set of correlated sensors is identified based on the following criteria: (1) a calculated similarity score between the sensor data from the first sensor and the sensor data from the sensors within the set of related sensors; (2) the geolocation of the first sensor and the sensors within the set of related sensors; and / or (3) one or more time series of the sensors within the first sensor and the sensors within the set of related sensors.

[0073] As described above, this disclosure introduces several data-driven approaches for automatically detecting, predicting, and repairing sensor failures in real time. Fault detection, prediction, repair, and tolerance are provided as needed (e.g., just-in-time models) and can be applied in real time to a set of critical sensors while avoiding unnecessary inspections and unnecessary monitoring or maintenance of non-critical sensors. Both physical sensors and / or virtual sensors (from digital twin models) are incorporated into this solution framework, and virtual sensors provide fault tolerance to physical sensors in some aspects. This disclosure enables operators or technicians to distinguish between sensor-versus-system component failures. In addition, fault tolerance from similar sensors already installed can be identified and utilized to reduce the cost of introducing new sensors to enable fault tolerance or the costs associated with critical sensor failures. Furthermore, sensor failures can be predicted, making it possible to perform or schedule repairs at a more convenient or less costly time (e.g., before incurring costs associated with sensor failures). The analysis may also be used to proactively introduce systematic fault tolerance, which is presented before any fault is detected.

[0074] Figure 9 shows an exemplary computing environment having exemplary computer equipment suitable for use in several exemplary implementations. The computer equipment 905 within the computing environment 900 may include one or more processing units, cores, or processors 910, memory 915 (e.g., RAM, ROM, and / or similar), internal storage 920 (e.g., magnetic, optical, solid-state storage, and / or organic) and / or an I / O interface 925, any of which may be connected on a communication mechanism or bus 930 for transmitting information, or may be incorporated within the computer equipment 905. Depending on the desired implementation, the I / O interface 925 may also be configured to receive images from a camera or provide images to a projector or display.

[0075] Computer device 905 may be communicatively connected to an input / user interface 935 and an output device / interface 940. One or both of the input / user interface 935 and the output device / interface 940 may be wired or wireless interfaces and may be detachable. The input / user interface 935 may include any physical or virtual device, component, sensor, or interface that can be used to provide input (e.g., buttons, touchscreen interfaces, keyboards, pointing / cursor controls, microphones, cameras, Brailles, motion sensors, accelerometers, optical readers, and / or similar). The output device / interface 940 may include displays, televisions, monitors, printers, speakers, Brailles, or similar. In some exemplary implementations, the input / user interface 935 and the output device / interface 940 may be integrated with or physically connected to the computer device 905. In other exemplary implementations, other computer devices may function as or provide input / user interfaces 935 and output devices / interfaces 940 for computer device 905.

[0076] Examples of computer devices 905 may include, but are not limited to, highly mobile devices (e.g., smartphones, devices in vehicles or other machines, devices carried by humans and animals, and similar devices), mobile devices (e.g., tablets, notebooks, laptops, personal computers, portable televisions, radios, and similar devices), and devices not designed for mobility (e.g., desktop computers, other computers, information kiosks, televisions, radios, and similar devices having one or more processors built in and / or connected thereto).

[0077] Computer device 905 may be communicably connected (for example, via I / O interface 925) to external storage 945 and network 950 for communication with any number of network-connected components, devices, and systems, including one or more computer devices of the same or different configurations. Computer device 905 or any connected computer device may function as, provide, or be referred to as a server, client, thin server, general-purpose machine, special-purpose machine, or other label.

[0078] The IO interface 925 may include, but is not limited to, wired and / or wireless interfaces that use any communication or IO protocol or standard (e.g., Ethernet, 902.11x, Universal Serial Bus, WiMAX, modem, cellular network protocol, and similar) to communicate information to and from at least all connected components, devices, and networks within the computing environment 900. The network 950 may be any network or combination of networks (e.g., the Internet, local area network, wide area network, telephone network, cellular network, satellite network, and similar).

[0079] The computer device 905 may use and / or communicate using computer-usable or computer-readable media, including temporary and non-temporary media. Temporary media include transmission media (e.g., metal cables, optical fibers), signals, carrier waves, and similar entities. Non-temporary media include magnetic media (e.g., disks and tapes), optical media (e.g., CD-ROMs, digital video discs, Blu-ray discs), solid-state media (e.g., RAM, ROMs, flash memory, solid-state storage), and other non-volatile storage or memory.

[0080] Computer device 905 may be used to implement techniques, methods, applications, processes, or computer executable instructions in several exemplary computing environments. Computer executable instructions may be obtained from temporary media and stored on and retrieved from non-temporary media. Executable instructions may originate from one or more arbitrary programming, scripting, and machine languages ​​(e.g., C, C++, C#, Java, Visual Basic, Python, Perl, JavaScript, and others).

[0081] The processor 910 may run under any operating system (OS) (not shown) in a native or virtual environment. One or more applications may be deployed with the OS and other applications (not shown), including a logic unit 960, an application programming interface (API) unit 965, an input unit 970, an output unit 975, and an inter-unit communication mechanism 995 for different units to communicate with each other. The units and elements described may be modified in design, function, configuration, or implementation, and are not limited to the description provided. The processor 910 may be in the form of a hardware processor such as a central processing unit (CPU) or a combination of hardware and software units.

[0082] In some exemplary implementations, information or execution instructions, upon receipt by the API unit 965, may be transmitted to one or more other units (e.g., logic unit 960, input unit 970, output unit 975). In some cases, the logic unit 90 may be configured to control the flow of information between units and to control the services provided by the API unit 965, input unit 970, and output unit 975 in some exemplary implementations described above. For example, the flow of one or more processes or implementations may be controlled by the logic unit 960 alone or in conjunction with the API unit 965. The input unit 970 may be configured to take input for a computation described in an exemplary implementation, and the output unit 975 may be configured to provide output based on a computation described in an exemplary implementation.

[0083] The processor 910 may be configured to receive sensor data from multiple related sensors. The processor 910 may also be configured to identify a set of correlated sensors in the multiple related sensors for a first sensor in the multiple related sensors. The processor 910 may be further configured to detect a fault in the first sensor based on at least one of the sensor data received from the sensor, the sensor data received from the set of correlated sensors, and the sensor data received from other sensors. The processor 910 may be further configured to implement a repair strategy based on the detected fault in the first sensor. The processor 910 may be further configured to calculate a similarity score between the sensor data from the first sensor and the sensor data from the sensors in the multiple related sensors. The processor 910 may also be configured to identify a set of correlated sensors based on the similarity score between the sensor data from the first sensor and the sensor data from the set of correlated sensors exceeding a threshold similarity score. The processor 910 may be configured to calculate a macro-similarity score based on the full set of time-series sensor data from the first sensor during a first period and the full set of time-series data from each sensor in the multiple related sensors during a first period. The processor 910 may also be configured to calculate multiple microsimilarity scores based on multiple subsets of the full set of time-series data from a first sensor during a first period and multiple subsets of the full set of time-series data from each sensor at multiple related sensors during the first period. The processor 910 may also be configured to calculate (or predict) multiple anomaly scores, including at least two of univariate anomaly scores, bivariate anomaly scores, or multivariate anomaly scores. The processor 910 may also be configured to predict a failure score indicating the likelihood of sensor failure by generating an ensemble of anomaly scores based on the multiple anomaly scores.

[0084] Some parts of the detailed explanation have been presented concerning algorithms and symbolic representations of operations within a computer. These algorithmic descriptions and symbolic representations are means used by those skilled in data processing technology to convey the essence of the innovation to others skilled in the art. An algorithm is a set of predefined steps that produce a desired end state or result. In one implementation example, the steps performed require the physical manipulation of tangible quantities to achieve a specific result.

[0085] Unless otherwise specified, explanations that use terms such as “processing,” “computing,” “calculating,” “determining,” and “displaying” throughout the explanation, as is evident from the explanation, should be understood to include actions and processes of a computer system or other information processing device that manipulate data represented as physical (electronic) quantities in the registers and memory of a computer system and convert it into other data similarly represented as physical quantities in the memory or registers of a computer system, or in other information storage, transmission, or display devices.

[0086] Implementation examples may also relate to apparatus for performing the operations described herein. Such apparatus may be specifically constructed for a required purpose and may include one or more general-purpose computers that are selectively activated or reconfigured by one or more computer programs. Such computer programs may be stored in computer-readable media such as computer-readable storage media or computer-readable signal media. Computer-readable storage media may include, but are not limited to, tangible media such as optical disks, magnetic disks, read-only memory, random-access memory, solid-state devices and drives, or any other type of tangible or non-temporary media suitable for storing electronic information. Computer-readable signal media may include media such as carrier waves. The algorithms and representations presented herein are not specific to any particular computer or other apparatus. Computer programs may include pure software implementations containing instructions for performing the operations in a desired implementation form.

[0087] Various general-purpose systems may be used with the programs and modules illustrated herein, or it may be convenient to construct more specialized devices for performing desired method steps. Furthermore, the implementation examples do not describe any particular programming language. It will be understood that various programming languages ​​may be used to implement the teachings of the implementation examples described herein. Instructions in a programming language may be executed by one or more processing units, such as a central processing unit (CPU), processor, or controller.

[0088] As is known in the art, the operations described above can be performed by hardware, software, or any combination of software and hardware. Various aspects of the implementation examples may be implemented using circuits and logic devices (hardware), while other aspects may be implemented using instructions (software) stored on a machine-readable medium, which, when executed by a processor, cause the processor to execute a method for performing the implementation of the present application. Furthermore, some implementation examples of the present application may be performed by hardware alone, while others may be performed by software alone. Moreover, the various functions described may be performed within a single unit or distributed across several components in any number of ways. When performed by software, the method may be executed by a processor such as a general-purpose computer based on instructions stored on a computer-readable medium. If desired, the instructions may be stored on the medium in compressed and / or encrypted form.

[0089] Furthermore, other implementations of the Application will become apparent to those skilled in the art by examining this Specification and practicing the Techniques of the Application. Various aspects and / or components of the implementations described herein may be used individually or in any combination. This Specification and the implementations are intended to be considered merely as examples, and the true scope and spirit of the Application are indicated by the appended claims.

Claims

1. Receiving sensor data from multiple related sensors, With respect to the first sensor in the plurality of related sensors, the set of correlation sensors in the plurality of related sensors is identified, A fault in the first sensor is detected based on at least one of the sensor data received from the first sensor, the sensor data received from the set of correlation sensors, and the sensor data received from other sensors. A repair strategy is implemented based on the fault detected by the first sensor, Includes, Detecting the failure in the first sensor includes detecting a predicted failure in the first sensor, and detecting the predicted failure in the first sensor is based on a sequence prediction model based on a deep learning recurrent neural network, the sequence prediction model is To predict multiple anomaly scores, including at least two of the following: univariate anomaly scores, bivariate anomaly scores, or multivariate anomaly scores. By generating an ensemble of abnormality scores based on the aforementioned multiple abnormality scores, a predictive failure score indicating the likelihood of failure in the first sensor is calculated. A method that includes this.

2. The method according to claim 1, wherein the plurality of related sensors are sensors that monitor the same system.

3. The method according to claim 1, wherein the plurality of associated sensors include at least one of physical sensors installed in the system or virtual sensors derived from a set of physical sensors based on a physical-based model.

4. The method according to claim 1, wherein the first sensor is a first critical sensor, the critical sensor is a sensor that captures critical data for monitoring the health of an underlying system and is used in at least one of identifying a remediation strategy, deriving business insights, or constructing solutions for a set of downstream tasks.

5. The method according to claim 1, wherein the set of correlation sensors includes a set of sensors having outputs that correlate with the output of the first sensor.

6. Receiving sensor data from multiple related sensors, With respect to the first sensor in the plurality of related sensors, the set of correlation sensors in the plurality of related sensors is identified, A fault in the first sensor is detected based on at least one of the sensor data received from the first sensor, the sensor data received from the set of correlation sensors, and the sensor data received from other sensors. A repair strategy is implemented based on the fault detected by the first sensor, Includes, The set of correlation sensors includes a plurality of sensors in the plurality of related sensors, the output of at least one of the plurality of sensors is not correlated with the output of the first sensor, and identifying the set of correlation sensors is a method based on a function of the outputs of the plurality of sensors that correlate with the first sensor.

7. Identifying the set of correlation sensors means that The process involves calculating a similarity score between the sensor data from the first sensor and the sensor data from the multiple related sensors. Identifying the set of correlation sensors based on the similarity score between the sensor data from the first sensor and the sensor data from the set of correlation sensors exceeding a threshold similarity score, The method according to claim 1, including the method described in claim 1.

8. Receiving sensor data from multiple related sensors, With respect to the first sensor in the plurality of related sensors, the set of correlation sensors in the plurality of related sensors is identified, A fault in the first sensor is detected based on at least one of the sensor data received from the first sensor, the sensor data received from the set of correlation sensors, and the sensor data received from other sensors. A repair strategy is implemented based on the fault detected by the first sensor, Includes, Identifying the set of correlation sensors means that The process involves calculating a similarity score between the sensor data from the first sensor and the sensor data from the multiple related sensors. Identifying the set of correlation sensors based on the similarity score between the sensor data from the first sensor and the sensor data from the set of correlation sensors exceeding a threshold similarity score, Includes, Calculating the similarity score between the sensor data from the first sensor and the sensor data from the multiple related sensors is: Calculate a macro-similarity score based on the full set of time-series sensor data from the first sensor during the first period and the full set of time-series data from each of the multiple related sensors during the first period, or A plurality of micro-similarity scores are calculated based on a plurality of subsets of the full set of time-series data from the first sensor during the first period and a plurality of subsets of the full set of time-series data from each of the plurality of related sensors during the first period. A method that includes at least one of the following.

9. Receiving sensor data from multiple related sensors, With respect to the first sensor in the plurality of related sensors, the set of correlation sensors in the plurality of related sensors is identified, A fault in the first sensor is detected based on at least one of the sensor data received from the first sensor, the sensor data received from the set of correlation sensors, and the sensor data received from other sensors. A repair strategy is implemented based on the fault detected by the first sensor, Includes, Detecting a fault in the sensor based on at least one of the sensor data received from the sensor, the sensor data received from the set of correlation sensors, and the sensor data received from other sensors includes one or more models for detecting the fault in the sensor, and the one or more models are A univariate anomaly detection model based on the sensor data received from the first sensor, A bivariate anomaly detection model based on the sensor data received from the first sensor and the sensor data received from the set of correlation sensors, or A multivariate anomaly detection model based on sensor data received from a plurality of related sensors, including the sensor data received from the first sensor, the sensor data received from the set of correlated sensors, and the sensor data received from other sensors. Methods that include...

10. The method further includes using one or more ensemble models to detect the fault in the sensor, wherein the one or more ensemble models are An ensemble anomaly detection model based on the univariate anomaly detection model and the bivariate anomaly detection model, or Ensemble anomaly detection model based on the univariate anomaly detection model and the multivariate anomaly detection model. The method according to claim 9, including the method described in claim 9.

11. The method according to claim 1, wherein the repair strategy includes replacing the sensor data from the first sensor with sensor data from one or more sensors in the set of correlated sensors.

12. Receiving sensor data from multiple related sensors, With respect to the first sensor in the plurality of related sensors, the set of correlation sensors in the plurality of related sensors is identified, A fault in the first sensor is detected based on at least one of the sensor data received from the first sensor, the sensor data received from the set of correlation sensors, and the sensor data received from other sensors. A repair strategy is implemented based on the fault detected by the first sensor, Includes, The repair strategy includes replacing the sensor data from the first sensor with sensor data from one or more sensors in the set of correlated sensors, The repair strategy is a method based on a root cause analysis of the detected failure using one or more explainable artificial intelligence (AI) technologies.

13. The sensor data from one or more sensors in the set of correlation sensors is A calculated similarity score between the sensor data from the first sensor and the sensor data from the sensors in the set of associated sensors, The geolocation of the sensor in the set of the first sensor and the related sensor, or The time series of the sensors in the set of the first sensor and the related sensors The method according to claim 11, specified based on one or more of the following.

14. Memory and At least one processor connected to the memory and A device including, wherein, based at least in part on the information stored in the memory, the at least one processor, Receive sensor data from multiple related sensors, With respect to the first sensor in the plurality of related sensors, the set of correlation sensors in the plurality of related sensors is identified. Based on the sensor data received from the first sensor, the sensor data received from the set of correlation sensors, and at least one of the sensor data received from other sensors, a fault in the first sensor is detected. The system is configured to implement a repair strategy based on the detected fault of the first sensor, The aforementioned at least one processor is The similarity score between the sensor data from the first sensor and the sensor data from the multiple related sensors is determined as follows: Calculate a macro-similarity score based on the full set of time-series sensor data from the first sensor during the first period and the full set of time-series data from each of the multiple related sensors during the first period, or The calculation is performed by one of the following methods: calculating a plurality of microsimilarity scores based on a plurality of subsets of the full set of time-series data from the first sensor during the first period and a plurality of subsets of the full set of time-series data from each sensor in the plurality of related sensors during the first period; The set of correlation sensors is identified based on the similarity score between the sensor data from the first sensor and the sensor data from the set of correlation sensors exceeding a threshold similarity score. A device configured to identify the set of correlation sensors.

15. The failure in the first sensor includes a predicted failure of the first sensor, and the at least one processor performs a sequence prediction based on a deep learning recurrent neural network. To predict multiple anomaly scores, including at least two of the following: univariate anomaly scores, bivariate anomaly scores, or multivariate anomaly scores, and By generating an ensemble of abnormality scores based on the aforementioned multiple abnormality scores, a predictive failure score indicating the likelihood of failure in the first sensor is calculated. The apparatus according to claim 14, configured to detect the predicted failure of the first sensor.

16. A computer-readable medium in a device for storing computer executable code, wherein the code, when executed by a processor, is transmitted to the processor. Sensor data is received from multiple related sensors. With respect to the first sensor in the plurality of related sensors, the set of correlation sensors in the plurality of related sensors is identified. Based on the sensor data received from the first sensor, the sensor data received from the set of correlation sensors, and at least one of the sensor data received from other sensors, a fault in the first sensor is detected. Based on the fault detected by the first sensor, a repair strategy is implemented. When the aforementioned code is executed by the processor, the processor will: Based on the sensor data received from the first sensor, the sensor data received from the set of correlation sensors, and at least one of the sensor data received from other sensors, one or more models for detecting a fault in the sensor are used to detect the fault in the first sensor, and the one or more models are A univariate anomaly detection model based on the sensor data received from the first sensor, A bivariate anomaly detection model based on the sensor data received from the first sensor and the sensor data received from the set of correlation sensors, A multivariate anomaly detection model based on sensor data received from a plurality of related sensors, including sensor data received from the first sensor, sensor data received from the set of correlation sensors, and sensor data received from other sensors. An ensemble anomaly detection model based on the univariate anomaly detection model and the bivariate anomaly detection model, or Ensemble anomaly detection model based on the univariate anomaly detection model and the multivariate anomaly detection model. Computer-readable media, including [specific examples of computer-readable media].

17. A computer-readable medium in a device for storing computer executable code, wherein the code, when executed by a processor, is transmitted to the processor. Sensor data is received from multiple related sensors. With respect to the first sensor in the plurality of related sensors, the set of correlation sensors in the plurality of related sensors is identified. Based on the sensor data received from the first sensor, the sensor data received from the set of correlation sensors, and at least one of the sensor data received from other sensors, a fault in the first sensor is detected. Based on the fault detected by the first sensor, a repair strategy is implemented. The failure in the first sensor includes a predicted failure of the first sensor, and the code, when executed by the processor, tells the processor: Based on a sequence prediction model based on a deep learning recurrent neural network, To predict multiple anomaly scores, including at least two of the following: univariate anomaly scores, bivariate anomaly scores, or multivariate anomaly scores. A computer-readable medium that generates an ensemble of abnormality scores based on the plurality of abnormality scores, thereby calculating a predicted failure score indicating the likelihood of failure in the first sensor, and detecting the predicted failure of the first sensor.

Citation Information

Patent Citations

  • Anomaly detection system and method

    JP2016201088A

  • Systems and methods for automated data science processes

    JP2023537766A

  • Sensor fault detection and diagnosis for autonomous systems

    US20160217627A1

  • Systems and methods for an automated data science process

    WO2022039748A1