Root cause analysis using Granger causality

By conducting conditional independence tests on polynomial quantities through a greedy hill-climbing process and utilizing Granger causality analysis, the problem of difficulty in identifying the causes of mechanical system failures was solved, achieving efficient and autonomous identification of failure causes and occurrences, and promoting preventive maintenance.

CN114846449BActive Publication Date: 2025-12-02INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080086820.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-11
Filing Date
2020-12-04
Publication Date
2025-12-02
Estimated Expiration
2040-12-04

AI Technical Summary

Technical Problem

In large-scale industrial systems, the root cause of mechanical failures is difficult to identify accurately, making preventative maintenance of the system difficult. Furthermore, existing technologies often rely on expert analysis, which cannot identify the time of failure in a timely manner.

Method used

A greedy hill-climbing process is used to perform conditional independence tests with a polynomial number of parameters. Granger causality analysis is then used to identify the causes and occurrences of mechanical system failures, thereby reducing the number of conditional independence tests and improving efficiency and accuracy.

Benefits of technology

It achieves efficient and autonomous identification of the causes of mechanical system failures, reduces the running time and error rate of causal discovery algorithms, enables timely identification of failure occurrences, and promotes preventive maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114846449B_ABST
    Figure CN114846449B_ABST
Patent Text Reader

Abstract

A system includes: a memory (116) for storing computer-executable components; and a processor (120) operatively coupled to the memory (116) and capable of executing the computer-executable components stored in the memory (116). The computer-executable components may include a maintenance component (108) that can detect causes of mechanical system failures by employing a greedy hill-climbing process to perform a polynomial number of conditional independence tests.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to one or more root cause analyses that can be based on Granger causality between time series data variables, and more specifically, to determining the causes and / or onsets of one or more failures of a mechanical system based on one or more Granger causality between time series data variables. Background Technology

[0002] Mechanical failures in large-scale industrial systems can lead to significant economic losses. In many cases, a single component or set of components in the system may fail, and the failure then propagates to other components and / or parts of the system. This propagation of failure can result in a collective failure of the entire system. The one or more parts that initially fail are often referred to as the root cause of the system failure. Typically, there can be a delay between the root cause event and the collective system failure. For example, a failure may propagate from the root cause to the system failure over a period of time (e.g., days, weeks, and / or months). Routine analysis of system failures involves subject matter experts who can repair the system after the failure event. In some cases, experts can identify the root cause of the system failure; however, it is often impossible for experts to accurately identify when the root cause began (e.g., the onset of the system failure). Therefore, there is a need in the art to address the aforementioned problems. Summary of the Invention

[0003] The following provides an overview to give a basic understanding of one or more embodiments of the present invention. This overview is not intended to identify key or essential elements, or to depict any scope of a particular embodiment or any scope of the claims. Its sole purpose is to present concepts in a simplified form as a preface to the specific implementations presented later. In one or more embodiments described herein, systems, computer-implemented methods, apparatuses, and / or computer program products are described that can consider root cause analysis based on one or more Granger causal relationships between time-series data variables.

[0004] From a first aspect, the present invention provides a system comprising: a memory storing computer-executable components; and a processor operatively coupled to the memory and executing the computer-executable components stored in the memory, wherein the computer-executable components include: a maintenance component that detects the causes of failures in a mechanical system by employing a greedy hillclimbing process to perform a polynomial number of conditional independence tests to determine Granger causality between variables from time-series data of a mechanical system under a given set of conditions.

[0005] In another aspect, the present invention provides a system comprising: a memory storing computer-executable components; and a processor operatively coupled to the memory and executing the computer-executable components stored in the memory, wherein the computer-executable components include: a maintenance component that detects the onset of mechanical system failures by employing a greedy hill-climbing process to perform a polynomial number of conditional independence tests to determine Granger causality between variables from time-series data of a mechanical system under a given set of regulation.

[0006] In another aspect, the present invention provides a computer-implemented method comprising: a system operatively coupled to a processor performing a polynomial number of conditional independence tests by employing a greedy hill-climbing process to determine Granger causality between variables from time-series data of a mechanical system under a given set of conditions, thereby detecting the cause of failure in the mechanical system.

[0007] In another aspect, the present invention provides a computer-implemented method comprising: a system operatively coupled to a processor detecting the onset of a mechanical system failure by employing a greedy hill-climbing process to perform a polynomial number of conditional independence tests to determine Granger causality between variables from time-series data of a mechanical system under a given set of conditions.

[0008] In another aspect, the present invention provides a computer program product for identifying information about faults in a mechanical system, the computer program product including a computer-readable storage medium that can be read by processing circuitry and stores instructions that are executed by the processing circuitry to perform a method for carrying out the steps of the present invention.

[0009] In another respect, the present invention provides a computer program stored on a computer-readable medium and loadable into the internal memory of a digital computer, comprising a software code portion that, when the program is run on a computer, performs the steps of the present invention.

[0010] In another aspect, the present invention provides a computer program product for analyzing faults in a mechanical system, the computer program product comprising a computer-readable storage medium having program instructions embodied therein, the program instructions being executable by a processor to enable the processor to determine Granger causality between variables from time-series data of a mechanical system under a given set of conditions by employing a greedy hill-climbing process to perform a polynomial number of conditional independence tests, thereby detecting fault occurrences in the mechanical system.

[0011] According to one embodiment, a system is provided. The system may include a memory capable of storing computer-executable components. The system may also include a processor operatively coupled to the memory and capable of executing the computer-executable components stored in the memory. The computer-executable components may include a maintenance component that can detect causes of failure in a mechanical system by employing a greedy hill-climbing process to perform a polynomial number of conditional independence tests to determine Granger causality between variables from time-series data of the mechanical system under a given set of conditions. An advantage of such a system is that the efficiency of executing causal discovery algorithms can be improved due to the use of a polynomial number of conditional independence tests instead of an exponential number of tests.

[0012] According to another embodiment, a system is provided. The system may include a memory capable of storing computer-executable components. The system may also include a processor operatively coupled to the memory and capable of executing the computer-executable components stored in the memory. The computer-executable components may include a maintenance component that detects the onset of mechanical system failures by employing a greedy hill-climbing process to perform a polynomial number of conditional independence tests to determine Granger causality between variables from time-series data of the mechanical system under a given set of conditions. An advantage of such a system is that it can determine when the root cause initially triggers a deterioration in the operation of the mechanical system.

[0013] According to an embodiment, a computer-implemented method is provided. This computer-implemented method may include a system operatively coupled to a processor determining Granger causality between variables from time-series data of a mechanical system under a given set of conditions by employing a greedy hill-climbing process to perform a polynomial number of conditional independence tests, thereby detecting the cause of failure in the mechanical system. An advantage of this computer-implemented method is that it can use machine learning to minimize the need for human interaction to identify root causes.

[0014] According to another embodiment, a computer-implemented method is provided. This computer-implemented method may include a system operatively coupled to a processor detecting the onset of a mechanical system failure by employing a greedy hill-climbing process to perform a multinomial number of conditional independence tests to determine Granger causality between variables from time-series data of a mechanical system under a given set of conditions. An advantage of this computer-implemented method is that it reduces the runtime of the causal discovery algorithm used to perform the greedy hill-climbing process due to the use of a multinomial number of conditional independence tests.

[0015] According to one embodiment, a computer program product for analyzing faults in a mechanical system is provided. The computer program product may include a computer-readable storage medium having program instructions. The program instructions are executable by a processor to cause the processor to detect fault occurrences in the mechanical system by employing a greedy hill-climbing process to perform a polynomial number of conditional independence tests, thereby determining Granger causality between variables from time-series data of the mechanical system under a given set of conditions. An advantage of this computer program product is that it can reduce the false discovery rate of causal discovery algorithms used to perform the greedy hill-climbing process. Attached Figure Description

[0016] This patent or application document contains at least one color drawing. A copy of this patent or application disclosure with one or more color drawings will be provided by the official upon request and payment of the necessary fees.

[0017] Figure 1 A block diagram of an example non-limiting system is shown, according to one or more embodiments described herein, which can determine the root cause of mechanical system failure and / or the onset of mechanical system failure to facilitate preventive maintenance.

[0018] Figure 2 A block diagram of an example non-limiting system is shown, which can construct one or more Granger causal relationship structures about time series data to facilitate root cause analysis according to one or more embodiments described herein;

[0019] Figure 3 A diagram illustrating an example non-restricted causal discovery algorithm that can be used to determine one or more Granger causal relationships between variables in time series data according to one or more embodiments described herein;

[0020] Figure 4 A block diagram of an example non-limiting system is shown, which can generate one or more adjacency matrices for time series data to facilitate root cause analysis according to one or more embodiments described herein;

[0021] Figure 5 A block diagram of an example non-limiting system is shown, which can cluster one or more adjacency matrices into two distinct cluster groups according to one or more embodiments described herein to facilitate root cause analysis.

[0022] Figure 6 A block diagram of an example non-limiting system for identifying the onset of faults in a clustered adjacency matrix system that can be based on Granger causality between variables characterizing time series data, according to one or more embodiments described herein;

[0023] Figure 7A block diagram of an example non-limiting system for determining one or more root causes and / or the onset of a system failure based on time series data according to one or more embodiments described herein;

[0024] Figure 8 The illustration shows an example non-limiting graph illustrating the effectiveness of one or more autonomous root cause analyses based on one or more Granger causal relationships found in time series data, according to one or more embodiments described herein.

[0025] Figure 9A The illustration shows an example non-limiting graph illustrating the effectiveness of one or more autonomous root cause analyses based on one or more Granger causal relationships found in time series data, according to one or more embodiments described herein.

[0026] Figure 9B The illustration shows an example non-limiting graph illustrating the effectiveness of one or more autonomous root cause analyses based on one or more Granger causal relationships found in time series data, according to one or more embodiments described herein.

[0027] Figure 9C The illustration shows an example non-limiting graph illustrating the effectiveness of one or more autonomous root cause analyses based on one or more Granger causal relationships found in time series data, according to one or more embodiments described herein.

[0028] Figure 10 The illustration shows an example non-limiting graph illustrating the effectiveness of one or more autonomous root cause analyses based on one or more Granger causal relationships found in time series data, according to one or more embodiments described herein.

[0029] Figure 11 A flowchart illustrating an exemplary non-limiting computer-implemented method that facilitates the analysis of one or more root causes of one or more occurrences of mechanical system failures according to one or more embodiments described herein;

[0030] Figure 12 A flowchart illustrating an example non-limiting computer implementation of a method that facilitates one or more root cause analyses of failures in one or more computer systems according to one or more embodiments described herein;

[0031] Figure 13 A flowchart illustrating an example non-limiting computer implementation of a method that facilitates one or more root cause analyses of failures in one or more computer systems according to one or more embodiments described herein;

[0032] Figure 14 A cloud computing environment according to one or more embodiments described herein is depicted;

[0033] Figure 15 An abstract model layer is described according to one or more embodiments described herein;

[0034] Figure 16 A block diagram illustrating an example non-limiting operating environment that may facilitate one or more embodiments described herein is shown. Detailed Implementation

[0035] The following detailed description is illustrative only and is not intended to limit the embodiments and / or their application or use. Furthermore, it is not intended to be construed as being bound by any express or implied information presented in the preceding Background or Summary of the Invention or Detailed Description sections.

[0036] One or more embodiments will now be described with reference to the accompanying drawings, wherein the same reference numerals are always used to denote the same elements. In the following description, numerous specific details are set forth for purposes of explanation in order to provide a more thorough understanding of one or more embodiments. However, it will be apparent in various circumstances that one or more embodiments may be practiced without these specific details.

[0037] Considering the problems with other implementations of root cause analysis for mechanical systems, this disclosure can be implemented to generate solutions to one or more of these problems by incorporating autonomous root cause analysis based on Granger causality found in time-series data, wherein autonomous root cause analysis can identify one or more root causes and / or occurrences of system failures. Advantageously, one or more embodiments described herein are capable of identifying when the root cause of a mechanical system failure occurs. Further, various embodiments described herein can identify the occurrence of system failures to facilitate preventative maintenance. Thus, the root cause analysis described in the various embodiments herein can identify the occurrence and / or root cause of system failures to enable preventative maintenance measures that can prevent system failures and the costs associated with system failures.

[0038] Various embodiments of the present invention may relate to computer processing systems, computer-implemented methods, apparatuses, and / or computer program products that facilitate efficient, effective, and autonomous (e.g., without direct human guidance) root cause analysis of mechanical system failures. For example, in one or more embodiments described herein, one or more greedy hill-climbing processes may be employed to perform a polynomial number of conditional independence tests to determine one or more Granger causal relationships between variables from time-series data of a mechanical system under a given set of conditions. One or more embodiments described herein may identify the onset of system failures based on Granger causality analysis. Furthermore, various embodiments described herein may identify one or more root causes of mechanical system failures and / or when one or more root causes occur based on Granger causality analysis.

[0039] Computer processing systems, computer-implemented methods, apparatuses, and / or computer program products employ hardware and / or software to solve inherently highly technical problems (e.g., root cause analysis of mechanical system failures) that are not abstract and cannot be performed by humans as a set of intellectual activities. For example, individuals or groups of individuals cannot readily monitor and / or analyze time-series data to determine Granger causality with the accuracy and efficiency of the various embodiments described herein. One or more embodiments described herein constitute a technical improvement over conventional techniques for discovering Granger causality between time-series data by reducing the number of conditional independence tests from an exponential number to a polynomial number. By reducing the number of conditional independence tests required for Granger causality discovery algorithms, one or more embodiments described herein can determine Granger causality at a lower statistical cost than that experienced by conventional techniques. Furthermore, the various embodiments described herein can focus on the practical application of greedy hill-climbing processes to perform a polynomial number of conditional independence tests to determine Granger causality between time-series data under a given set of conditions, thereby determining the onset and / or root cause of mechanical system failures.

[0040] Figure 1 A block diagram of an exemplary non-limiting system 100 capable of performing one or more root cause analyses is shown. For brevity, repeated descriptions of similar elements used in other embodiments described herein are omitted. Aspects of systems (e.g., system 100, etc.), apparatuses, or processes in various embodiments of the invention may constitute one or more machine-executable components embodied within one or more machines, for example, embodied in one or more computer-readable media (or media) associated with one or more machines. Such components, when executed by one or more machines (e.g., computers, computing devices, virtual machines, etc.), enable the machines to perform the described operations.

[0041] like Figure 1 As shown, system 100 may include one or more servers 102, one or more networks 104, and / or input devices 106. Server 102 may include a maintenance component 108. Maintenance component 108 may further include a communication component 110 and / or a partitioning component 112. Furthermore, server 102 may include at least one memory 116 or otherwise associated with it. Server 102 may also include a system bus 118, which may be coupled to various components, such as, but not limited to, maintenance component 108 and associated components, memory 116, and / or processor 120. Although in Figure 1 Server 102 is shown, but in other embodiments, multiple devices of various types can be connected. Figure 1 The features shown are associated with or include these features. Furthermore, server 102 can communicate with one or more cloud computing environments.

[0042] One or more networks 104 may include wired and wireless networks, including but not limited to cellular networks, wide area networks (WANs) (e.g., the Internet), or local area networks (LANs). For example, server 102 may communicate with one or more input devices 106 using virtually any desired wired or wireless technology (and vice versa), including, but not limited to, cellular, WAN, Wi-Fi, Wi-Max, WLAN, Bluetooth, and combinations thereof. Furthermore, although maintenance component 108 may be provided on one or more servers 102 in the illustrated embodiment, it should be understood that the architecture of system 100 is not limited thereto. For example, maintenance component 108 or one or more components of maintenance component 108 may reside at another computer device, such as another server device, client device, etc.

[0043] One or more input devices 106 may include one or more computerized devices, which may include, but are not limited to: personal computers, desktop computers, laptop computers, cellular phones (e.g., smartphones), computerized tablets (e.g., including processors), smartwatches, keyboards, touchscreens, mice, combinations thereof, etc. One or more input devices 106 may be used to input time-series data of a mechanical system into system 100, thereby sharing said data with server 102 (e.g., via a direct connection and / or via one or more networks 104). For example, one or more input devices 106 may transmit data to communication component 110 (e.g., via a direct connection and / or via one or more networks 104). Additionally, one or more input devices 106 may include one or more displays that can present one or more outputs generated by system 100 to a user. For example, one or more displays may include, but are not limited to: cathode ray tube displays (CRTs), light-emitting diode displays (LEDs), electroluminescent displays (ELDs), plasma display panels (PDPs), liquid crystal displays (LCDs), organic light-emitting diode displays (OLEDs), combinations thereof, and / or the like.

[0044] In various embodiments, one or more input devices 106 and / or one or more networks 104 may be used to input one or more settings and / or commands into system 100. For example, in the various embodiments described herein, one or more input devices 106 may be used to operate and / or manipulate server 102 and / or associated components. Additionally, one or more input devices 106 may be used to display one or more outputs (e.g., displays, data, visualizations, etc.) generated by server 102 and / or associated components. Further, in one or more embodiments, one or more input devices 106 may be included within and / or operatively coupled to a cloud computing environment.

[0045] In various embodiments, one or more input devices 106 may include one or more sensors and / or detectors included within a mechanical system, wherein the one or more input devices 106 may collect and / or measure time-series data. Examples of sensors that may be included within one or more input devices 106 may include, but are not limited to: temperature sensors, pressure sensors (e.g., barometer sensors, pressure gauge sensors, Burden tube pressure sensors, vacuum pressure sensors, combinations thereof, and / or the like), vibration sensors (e.g., accelerometers, strain gauges, capacitive displacement sensors, combinations thereof, and / or the like), ultrasonic sensors, touch sensors, proximity sensors, liquid level sensors, smoke sensors, gas sensors, valve measurement sensors, housing pressure measurement sensors, combinations thereof, and / or the like.

[0046] In one or more embodiments, one or more input devices 106 may be used to input time-series data into system 100 (e.g., via manual operation, autonomous collection, autonomous detection, and / or autonomous measurement). The time-series data may focus on the operation of one or more components included within a mechanical system. For example, the mechanical system may be included within an Internet of Things (IoT), where one or more input devices 106 may include one or more sensors and / or detectors that monitor the operation of one or more parts of the mechanical system. The time-series data may include data collected, detected, and / or measured by one or more input devices 106. For example, the time-series data may include data points describing the amount of vibration, temperature, and / or pressure experienced by one or more parts of the mechanical system during operation, and when the one or more parts experienced the data points.

[0047] For example, the operational characteristics of various parts of a mechanical system can be represented by one or more data points, which can be indexed in chronological order (e.g., indexed according to a sequence obtained at consecutive time points) to establish time-series data. For example, one or more input devices 106 can collect, detect, and / or measure the operational state of each part at given time intervals, wherein the data collected, detected, and / or measured at each time interval can be timestamped and added to the data collected, detected, and / or measured at previous time intervals to establish time-series data. Thus, time-series data can be input manually (e.g., via one or more users) or autonomously (e.g., via one or more sensors and / or detectors) via one or more input devices 106, and / or can characterize the operational state of various parts of the mechanical system over time.

[0048] In various embodiments, one or more input devices 106 may share time-series data with maintenance component 108 via communication component 110 and / or one or more networks 104. Communication component 110 may receive time-series data and share it with one or more associated components of maintenance component 108 as described herein (e.g., via system bus 118). In one or more embodiments, communication component 110 may receive time-series data and store it in one or more memories 116 for processing by maintenance component 108.

[0049] The segmentation component 112 can divide time-series data into multiple groups based on defined time intervals. Exemplary time intervals may include, but are not limited to, hours, days, weeks, months, years, and combinations thereof. In various embodiments, one or more input devices 106 may be used to define the time intervals. Additionally, the segmentation component 112 can store multiple data groups extracted from the time-series data in one or more memories 116. For example, when the time interval is one week, the segmentation component 112 can divide the time-series data into multiple groups, where each group includes time-series data associated with a given week. For example, when the time interval is one week, time-series data characterizing the operational status of one or more parts of a mechanical system during the first week of September may be divided into a first group; while time-series data characterizing the operational status of those one or more parts during the second week of September may be divided into a second group.

[0050] Figure 2 A diagram illustrating an example non-limiting system 100 further including a structural component 202 according to one or more embodiments described herein. For brevity, repeated descriptions of similar elements employed in other embodiments described herein are omitted. In various embodiments, the structural component may employ one or more causal discovery algorithms for each defined group of time-series data to generate inferred causal relationship structures for multiple groups.

[0051] In one or more embodiments, structural component 202 may employ one or more parametric or nonparametric conditional independence testers to infer causal structure among variables included within a time series data set. For example, structural component 202 may employ machine learning to perform one or more causal discovery algorithms that may use one or more conditional independence testers to perform tests to generate one or more Bayesian networks (e.g., directed graphs depicting relationships between variables in time series data) to determine the presence and / or absence of edges. Examples of conditional independence testers may include, but are not limited to, ParCorr, CCIT, and / or RCOT.

[0052] Structural component 202 may utilize one or more conditional independence testers to determine one or more p-values, which may correspond to the probability that the assumed relationship between variables included in the time series data is independent. For example, a p-value may represent the probability that a first time series data variable is independent of a second time series data variable given a third time series data variable. For example, a p-value may correspond to the probability that, given any time series data... Event I(X) i →X i ||X AThe probability that ) = 0; where "i" can correspond to a first time series data variable, "j" can correspond to a second time series data variable, "A" can correspond to a third time series data variable, and / or "m" can correspond to the number of features (e.g., the dimension of the feature space). In various embodiments, given a third time series data variable, a high p-value can indicate the independence between the first and second time series data variables given the third time series data variable; while a low p-value can indicate the dependency between the first and second time series data variables given the third time series data variable.

[0053] In one or more embodiments, structural component 202 may utilize one or more conditional independence testers to determine association measurements characterizing possible dependencies between time series data according to Equation 1 below.

[0054] Assoc a (i→j;A)=α-min(α,DI(i,j,A)) (1)

[0055] Here, "DI" can correspond to orientation information, and "a" can correspond to a defined p-value threshold. P-values ​​greater than the p-value threshold α can be clipped to the p-value threshold α. Therefore, the p-value threshold α can correspond to the maximum association between given time-series data variables. Thus, structural component 202 can determine the p-value using one or more conditional independence testers, and then correlate that p-value with an association measurement according to Equation 1.

[0056] Furthermore, structural component 202 can utilize association measures to perform a causal discovery algorithm on multiple time series data sets. During the first phase of the causal discovery algorithm, a causal graph (e.g., a Bayesian network) can be constructed based on p-values ​​and / or association measures facilitated by one or more conditional independence testers. Further, structural component 202 can refer to the causal graph to identify time series data variables that have a high association with the target time series data variable (e.g., exceeding a threshold set via one or more input devices 106) and include the identified time series data variables in a candidate parent set. Time series data variables can be repeatedly added to the candidate parent set until no more associations are found. For example, one or more conditional independence testers can perform a multinomial number of tests to find the variable with the greatest association with the target variable. During the second phase of the causal discovery algorithm, structural component 202 can prune the candidate parent set to remove one or more irrelevant variables (e.g., by removing spurious edges added to the causal graph during the first phase). Additionally, structural component 202 can incorporate one or more false discovery rate controls into the causal discovery algorithm.

[0057] Figure 3A diagram illustrates an exemplary non-limiting MMPC-p reduction algorithm 300, illustrative of one or more causal discovery algorithms executable by structural component 202, according to one or more embodiments described herein. For brevity, repeated descriptions of similar elements employed in other embodiments described herein are omitted. Figure 3 As shown, the MMPC-p reduction algorithm 300 can be a greedy hill-climbing algorithm based on Bayesian networks, which uses a conditional independence tester to identify a causal graph on a set of variables from time series data.

[0058] like Figure 3 As shown, the MMPC-p reduction algorithm 300 can begin by analyzing the candidate parents (CPs) of the time-series data variable j initialized at φ. Further, the MMPC-p reduction algorithm 300 can form a set of adjustment for the candidate parents based on the K value, which can be equal to the candidate parents minus the K value. In various embodiments, the K value can be set via one or more input devices 106, and / or the K value can be greater than or equal to zero. For example, in the case of 5 candidate parents and a K value of 1, the adjustment set can be equal to 4 candidates. Further, the number of conditional independence tests performed during the first phase of the MMPC-p reduction algorithm 300 can be a polynomial number of the K value. For example, the number n of conditional independence tests performed during the first phase of the MMPC-p reduction algorithm 300 can be equal to n0. K+1 Changing the value of K can be used to trade off the size of the conditioning set by the number of conditional independence tests performed. For example, a small value of K can result in a larger conditioning set, while a large value of K can result in more conditional independence tests being performed.

[0059] Among the candidates in the conditional set, the MMPC-p reduction algorithm 300 can find the minimum association between two given time series data variables i and j (e.g., via one or more conditional independence testers and / or the association measure defined by Equation 1). The MMPC-p reduction algorithm 300 then analyzes the parent candidates of time series data variable i that are not yet in the candidate parent set of time series data variable j, such that the association measure is maximized (e.g., conditioned on the current candidate parent set of j). If a candidate parent of i is already in the candidate parent set of j, the association measure will be zero, and therefore not the maximum association. When a candidate parent of i is found to have the maximum association measure (e.g., a value greater than zero), the candidate parent of i is added to the candidate parent set of j. Thus, candidate parents outside the initial candidate parent set of j can be analyzed to identify associations that would mark their inclusion in the candidate parent set of j (e.g., characterized by an association measure greater than zero).

[0060] Once the first phase of the MMPC-p reduction algorithm 300 has established the candidate parent set of j (e.g., as described above and in...), Figure 3 As shown in lines 1-7), the second stage of the MMPC-p reduction algorithm 300 can prune the candidate parent set (e.g., as shown in lines 1-7). Figure 3 (As shown in lines 8-20). For each member Y in the candidate parent set, the MMPC-p reduction algorithm 300 can determine the association between Y and j conditioned on the remaining members in the candidate parent set (e.g., via one or more conditional independence testers and / or the association measure characterized by Equation 1) to identify the member with the minimum association. If the minimum association is equal to zero, the given member can be removed from the candidate parent set. If the minimum association is greater than zero, the given member can be maintained. Thus, the MMPC-p reduction algorithm 300 can prune the composition of the candidate parent set.

[0061] In one or more embodiments, the second stage of the MMPC-p reduction algorithm 300 may further determine what p-value a given parent is added to the candidate parent set, the candidate parent set being determined by... Figure 3 In lines 1, 19, 21 and / or 22 Indicate. For example, This can be used to determine one or more statistical guarantees regarding the edges of a causal graph. Additionally, in one or more embodiments, the structural component 202 can generate a causal relationship structure for each group of time-series data (e.g., by performing a causal discovery algorithm, such as the MMPC-p reduction algorithm 300).

[0062] Figure 4 A diagram is shown illustrating an example non-limiting system 100 further including a matrix component 402 according to one or more embodiments described herein. For brevity, repeated descriptions of similar elements used in other embodiments described herein are omitted. In various embodiments, the matrix component 402 may construct multiple adjacency matrices for each group of time-series data based on one or more causal structures derived from the structural component 202.

[0063] For example, matrix component 402 can construct an adjacency matrix for each causal graph constructed by structural component 202. The adjacency matrix can depict the parent of each time series data variable in a given group based on the determination of a causal discovery algorithm (e.g., MMPC-p reduction algorithm 300). For example, the adjacency matrix can be composed in Boolean format (e.g., each row can index various variables within the group; each column can index possible parents; as a parent candidate determined by the causal discovery algorithm can correspond to a value of 1 in the matrix; as a non-parent determined by the causal discovery algorithm can correspond to a value of 0 in the matrix).

[0064] Figure 5A diagram is shown illustrating an example non-limiting system 100 further including a clustering component 502 according to one or more embodiments described herein. For brevity, repeated descriptions of similar elements employed in other embodiments described herein are omitted. In various embodiments, the clustering component 502 may employ machine learning to perform one or more K-means clustering techniques to cluster multiple adjacency matrices into different cluster groups.

[0065] For example, a first cluster group may include an adjacency matrix characterizing the standard operation of the mechanical system. For example, the first cluster group may include an adjacency matrix that includes Granger causality between time series data during time intervals in which the mechanical system does not experience a fault (e.g., each part monitored by sensors of one or more input devices 106 operates in a standard manner, such as within expected tolerances). A second cluster group may include an adjacency matrix characterizing the non-standard operation of the mechanical system. For example, the second cluster group may include an adjacency matrix that includes Granger causality between time series data during time intervals in which the mechanical system is experiencing a fault (e.g., one or more parts monitored by sensors of one or more input devices 106 operate in a non-standard manner, such as outside expected tolerances).

[0066] In one or more embodiments, the data topology of the adjacency matrix may vary depending on whether a given part of the mechanical system has experienced a failure, based on time-series data describing the operation of that given part. Thus, the clustering component 502 may cluster the adjacency matrix into separate clusters based on the variation in data topology: clustering adjacency matrices associated with standard operations of the mechanical system into a first cluster group, and clustering adjacency matrices associated with non-standard operations (e.g., failures) of the mechanical system into a second cluster group.

[0067] Figure 6 A diagram is shown of an example non-limiting system 100 further including an outbreak component 602 according to one or more embodiments described herein. For brevity, repeated descriptions of similar elements used in other embodiments described herein are omitted. In various embodiments, the outbreak component 602 may analyze a clustered adjacency matrix to identify the outbreak of a root cause failure of a mechanical system.

[0068] For example, the outbreak component 602 can identify one or more adjacency matrices that are adjacent to one or more other adjacency matrices of another cluster group within the time series of time series data. For example, a transition between cluster groups across adjacency matrices arranged in the time series can mark a transition between standard and non-standard operation of a mechanical system, which can indicate a failure experienced by one or more parts represented by the time series data. For example, the outbreak component 602 can sort clustered adjacency matrices into a time series of time series data based on time intervals associated with a group of time series data represented by a given adjacency matrix. The outbreak component 602 can determine the outbreak of a root cause by identifying instances along the time series where the adjacency matrix of a first cluster group is adjacent to the adjacency matrix of a second cluster group. Thus, the outbreak component 602 can determine when a root cause occurs and / or the outbreak of a root cause based on data topology changes in the Granger causality of the time series data (such as those depicted by transitions along the time series from the adjacency matrix of the first cluster group to the adjacency matrix of the second cluster group).

[0069] Figure 7 A diagram illustrating an exemplary non-limiting system 100 further including a causation component 702 according to one or more embodiments described herein is shown. For brevity, repeated descriptions of similar elements used in other embodiments described herein are omitted. In various embodiments, the causation component 702 may analyze the adjacency matrix of identified cluster group transitions (e.g., as identified by the causation component 602) to identify one or more time-series data variables corresponding to the maximum change between the identified adjacency matrices. In one or more embodiments, the variable is associated with the maximum change between the identified adjacency matrices and indicates the root cause of a mechanical system failure, or a potential root cause of a mechanical system failure.

[0070] For example, the causation component 702 can determine the Hamming distance between time-series data variables in an adjacency matrix (e.g., an adjacency matrix defining cluster group transitions) identified by the outbreak component 602. For example, the causation component 702 can determine the Hamming distance between variables included in a first adjacency matrix and a second adjacency matrix; wherein the first adjacency matrix may belong to a first cluster group and is adjacent to a second adjacency matrix that may belong to a second cluster group. Based on the Hamming distance, the causation component 702 can identify the number of defined variables (e.g., five variables) associated with the maximum change between the two adjacency matrices. In one or more embodiments, the number of defined variables can be set via one or more input devices 106. One or more parts of the mechanical system characterized by the identified variables can be the root cause of past failures or the root cause of potential failures. Further, the time when the adjacency matrix transitions from one cluster group to another can be the time when the identified part (e.g., as described by the identified variables) experiences a failure (e.g., the time when the outbreak of the root cause occurs).

[0071] Figure 8 Figures 800 and 802, illustrating example non-limiting graphs 800 and 802, are shown according to one or more embodiments described herein, demonstrating the effectiveness of the MMPC-p reduction algorithm 300 using various conditional independence testers. For brevity, repeated descriptions of similar elements employed in other embodiments described herein are omitted. Figures 800 and 802 focus on performing causal discovery algorithms on a comprehensive dataset containing 10 time-series variables generated by a Kuramoto oscillator.

[0072] As shown in Table 800, line 804 corresponds to a traditional causal discovery algorithm that utilizes an exponential number of conditional independence tests to construct a Granger causal graph. For example, line 804 could correspond to performing the MMPC algorithm on time series data using the CCIT conditional independence tester and an α value of 0.1. Lines 806 and 808 focus on performing the Bayesian network-based greedy hill-climbing causal discovery algorithm described in this paper, which employs a multinomial number of conditional independence tests to construct a Granger causal graph. For example, line 806 could correspond to performing the MMPC-p reduction algorithm 300 using the CCIT conditional independence tester and an α value of 0.1. Similarly, line 808 could correspond to performing the MMPC-p reduction algorithm 300 using the ParCorr conditional independence tester and an α value of 0.1. As shown in Figures 800 and / or 802, compared with conventional techniques, implementing the Bayesian network-based greedy hill-climbing causal discovery algorithm (e.g., MMPC-p reduction algorithm 300) described in this paper can achieve improved false positive rate and / or false negative rate.

[0073] Figure 9A , 9B Figures 900, 902, and / or 904 illustrate exemplary non-limiting diagrams 900, 902, and / or 904 that demonstrate the effectiveness of root cause analysis that can be performed by maintenance component 108 according to various embodiments described herein. For brevity, repeated descriptions of similar elements employed in other embodiments described herein are omitted. Figures 900, 902, and / or 904 may focus on time-series data including sensor readings from a steam turbine. The unique operation of a steam turbine can result in one or more properties of the time-series data that can make causal discovery challenging. For example, the time-series data may exhibit time-based changing properties. For example, a steam turbine does not operate most nights. Moreover, during the day, a steam turbine may have various operating patterns that can alter the relationships between variables. In another example, the time-series data may characterize lag dependencies and high-frequency variations. The behavior of a variable can change rapidly; however, the realization of the effect of the change in a variable on other variables can be delayed. For example, the vibration of the rotor may change within seconds, but the effect on the power output of the steam turbine may take several minutes to observe.

[0074] To combat the challenges of causal discovery, maintenance component 108 can implement the following heuristics when exporting the analysis presented in graphs 900, 902, and / or 904. First, structure component 202 can perform a causal discovery algorithm (e.g., MMPC-p reduction algorithm 300) on five bootstraps of the time series dataset and construct a weighted causal graph. The weights of the edges in the causal graph can represent the number of bootstraps in which that edge was detected. Second, structure component 202 can subsample the time series data at two-minute intervals, allowing the causal discovery algorithm (e.g., MMPC-p reduction algorithm 300) to discover relationships regardless of the aforementioned lag.

[0075] Figure 900 depicts Granger causality of time-series data plotted along a time series for thrust vibrations (e.g., measured by one or more sensors of input device 106) at seven bearings within a steam turbine. Figure 9A As shown, the partitioning component 112 can organize time series data into multiple groups (e.g., Figure 9A Groups 0-7 are shown. Chart 902 depicts a magnified portion of Chart 900 from September 15th to September 29th. (As shown) Figure 9B As shown, clustering component 502 can cluster the adjacency matrices of multiple groups into two cluster groups (e.g., Figure 9B(As shown in cluster groups 1 and 2). For example, time series data groups 0, 1, and 2 may have similar data topologies and be clustered into cluster group 1; while time series data group 3 may have different data topologies than time series data groups 0, 1, and 2 and may be clustered into cluster group 2. Thus, the outbreak component 602 can identify cluster group transitions between the adjacency matrices associated with time series data groups 2 and 3 to determine the root cause of the outbreak.

[0076] Furthermore, the cause component 702 can determine the Hamming distance between the variables in the adjacency matrix associated with time series data group 2 and the adjacency matrix associated with time series data group 3. For example... Figure 9B As shown, the variables associated with bearings 1, 4, 6, and 2 can be identified (e.g., cause component 702) as having the maximum Hamming distance. Figure 904 depicts a magnified portion of Figure 900 from October 7th to October 19th. Figure 9C As shown, due to the outbreak identified between time series groups 2 and 3 (e.g., the outbreak on September 22), bearings 1, 4, 6, and 2 may begin to have causal effects on variables they did not previously cause. Further, the causal effect is shown as a visible alteration of operational behavior in time series data group 5; therefore, maintenance component 108 can identify the outbreak that occurred in time series data group 2 and manifested as a significant deterioration in time series data group 5, and can identify the possible variables causing the root cause that are associated with the operation of bearings 1, 4, 6, and 2.

[0077] Figure 10 Exemplary non-limiting diagram 1000 is shown, illustrating the effectiveness of root cause analysis performed by maintainable component 108 according to various embodiments described herein. For brevity, repeated descriptions of similar elements used in other embodiments described herein are omitted. Diagram 1000 may further focus on… Figure 9A , 9B And / or time-series data of the steam turbine analyzed in 9C. The steam turbine experienced a second system failure on July 6, and Figure 1000 depicts the onset of the root cause, as determined by maintenance component 108 according to the various embodiments described herein. As shown in Figure 1000, the adjacency matrix associated with time-series data group 7 and the adjacency matrix associated with time-series data group 8 can be clustered into different cluster groups, thereby indicating the onset of the root cause. Figure 1000 also shows some readings associated with pressure and / or control values ​​decreasing over time, where a failure in one or more pressure and / or control values ​​at a cluster group transition may correspond to the onset of the root cause.

[0078] Figure 11A flowchart of an example non-limiting computer-implemented method 1100 that facilitates one or more root cause analyses based on time series data, according to one or more embodiments described herein, is shown. For brevity, repeated descriptions of similar elements used in other embodiments described herein are omitted.

[0079] At 1102, the computer-implemented method 1100 may include receiving (e.g., via communication component 110 and / or input device 106) time-series data about a mechanical system by a system 100 operatively coupled to processor 120. For example, the time-series data may focus on the operational state of one or more parts included within the mechanical system. In various embodiments, the time-series data may be collected, determined, and / or measured by one or more sensors and / or detectors monitoring the operation of the mechanical system. Further, in one or more embodiments, the time-series data may be received via one or more networks 104 (e.g., utilizing one or more cloud computing environments).

[0080] At 1104, the computer-implemented method 1100 may include determining Granger causality between variables from time-series data under a given set of conditions by employing a greedy hill-climbing process to perform a polynomial number of conditional independence tests, and the system 100 may detect (e.g., via maintenance component 108) the cause of failure in the mechanical system. For example, according to the various embodiments described herein, the detection at 1104 may include (e.g., via causation component 602) identifying cluster group transitions between adjacency matrices characterizing Granger causality. Further, according to the various embodiments described herein, the detection at 1104 may include determining (e.g., via causation component 702) the Hamming distance between variables in the adjacency matrices defining the cluster group transitions. In one or more embodiments, the variable having the largest Hamming distance may be identified as associated with the root cause of the failure.

[0081] Figure 12 A flowchart illustrating an example non-limiting computer-implemented method 1200 that facilitates one or more root cause analyses based on time series data according to one or more embodiments described herein is shown. For brevity, repeated descriptions of similar elements used in other embodiments described herein are omitted.

[0082] At 1202, the computer-implemented method 1200 may include receiving (e.g., via communication component 110 and / or input device 106) time-series data about a mechanical system by a system 100 operatively coupled to processor 120. For example, the time-series data may focus on the operational state of one or more parts included within the mechanical system. In various embodiments, the time-series data may be collected, determined, and / or measured by one or more sensors and / or detectors monitoring the operation of the mechanical system. Further, in one or more embodiments, the time-series data may be received via one or more networks 104 (e.g., utilizing one or more cloud computing environments).

[0083] At 1204, the computer-implemented method 1200 may include determining Granger causality between variables from time-series data under a given set of conditions by employing a greedy hill-climbing process to perform a polynomial number of conditional independence tests, and the system 100 may detect (e.g., via maintenance component 108) the occurrence of a mechanical system failure. For example, according to various embodiments described herein, the detection at 1204 may include (e.g., via occurrence component 602) identifying cluster group transitions between adjacency matrices characterizing Granger causality. Cluster group transitions may define changes marked in the data topology of Granger causality and / or may indicate when a failure within the mechanical system initially occurred. In one or more embodiments, the computer-implemented method 1200 may detect occurrences of experienced mechanical system failures to facilitate repair. In one or more embodiments, the computer-implemented method 1200 may detect occurrences of potential system failures to facilitate preventative maintenance.

[0084] Figure 13 A flowchart of an example non-limiting computer-implemented method 1300 that facilitates one or more root cause analyses based on time series data, according to one or more embodiments described herein, is shown. For brevity, repeated descriptions of similar elements used in other embodiments described herein are omitted.

[0085] At 1302, the computer-implemented method 1300 may include receiving (e.g., via communication component 110 and / or input device 106) time-series data about a mechanical system by a system 100 operatively coupled to processor 120. For example, the time-series data may focus on the operational state of one or more parts included within the mechanical system. In various embodiments, the time-series data may be collected, determined, and / or measured by one or more sensors and / or detectors monitoring the operation of the mechanical system. Further, in one or more embodiments, the time-series data may be received via one or more networks 104 (e.g., utilizing one or more cloud computing environments).

[0086] At 1304, the computer-implemented method 1300 may include dividing time-series data (e.g., via partitioning component 112) into fixed-length groups by system 100. For example, according to various embodiments described herein, the fixed length may be based on defined time intervals. At 1306, the computer-implemented method 1300 may include system 100 performing (e.g., via structural component 202) one or more causal discovery algorithms for each group to construct one or more Granger causal graphs characterizing Granger causality in the time-series data. For example, according to various embodiments described herein, the one or more causal discovery algorithms are greedy hill-climbing algorithms based on Bayesian networks with conditional independence testers to identify Granger causal graphs on time-series data, such as the MMPC-p reduction algorithm 300.

[0087] At 1308, the computer-implemented method 1300 may include one or more adjacency matrices constructed by system 100 based on a Granger causal graph (e.g., via matrix component 402). For example, according to various embodiments described herein, the one or more adjacency matrices may include candidate parents indexed for each variable in the time series data. At 1310, the computer-implemented method 1300 may include clustering the one or more adjacency matrices into two cluster groups by system 100 using K-means clustering (e.g., via clustering component 502). For example, according to various embodiments described herein, the first cluster group may be associated with one or more standard operating states of a mechanical system, and / or the second cluster group may be associated with one or more non-standard operating states of the mechanical system (e.g., operations including one or more malfunctions).

[0088] At 1312, the computer-implemented method 1300 may include adjacency matrices identified by system 100 (e.g., by the causation component 602) as adjacent to each other along a time series and belonging to different cluster groups. For example, according to the various embodiments described herein, the identified adjacency matrices may define cluster group transitions along a time series. Further, cluster group transitions may indicate the onset of a root cause and / or potential failure of the mechanical system. At 1314, the computer-implemented method 1300 may include Hamming distances between variables of the identified adjacency matrices determined by system 100 (e.g., via causation component 702). For example, a variable having the maximum Hamming distance between adjacency matrices defining cluster group transitions may be associated with the root cause of a failure or potential failure of the mechanical system.

[0089] It should be understood that although this disclosure includes a detailed description of cloud computing, the implementation of the teachings recorded herein is not limited to a cloud computing environment. Rather, embodiments of the invention can be implemented in conjunction with any other type of computing environment now known or developed hereafter.

[0090] Cloud computing is a service delivery model that enables convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing power, storage, applications, VMs, and services) that can be rapidly provisioned and released with minimal management costs or interaction with service providers. This cloud model may include at least five features, at least three service models, and at least four deployment models.

[0091] The characteristics are as follows:

[0092] On-demand self-service: Cloud consumers can unilaterally and automatically provide computing power (such as server time and network storage) on demand without human interaction with the service provider.

[0093] Wide network access: Capabilities are available on the network and accessed through standard mechanisms that facilitate the use of heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs).

[0094] Resource pooling: A provider's computing resources are grouped into resource pools to serve multiple consumers using a multi-tenant model, where different physical and virtual resources are dynamically allocated and reallocated based on demand. Typically, consumers cannot control or know the exact location of the resources provided, but can specify the location at a higher level of abstraction (e.g., country, state, or data center), thus exhibiting location independence.

[0095] Rapid flexibility: Capabilities can be rapidly and flexibly (in some cases automatically) provided to expand outward quickly and be rapidly released to shrink back down. For consumers, the available capacity often appears unlimited and can be purchased at any time and in any quantity.

[0096] Measurable services: Cloud systems automatically control and optimize resource usage by leveraging metering capabilities at a level of abstraction appropriate to the service type (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported, providing transparency to both service providers and consumers.

[0097] The service model is as follows:

[0098] Software as a Service (SaaS): The capability offered to consumers is the ability to use applications running on a provider's cloud infrastructure. These applications can be accessed from various client devices via thin client interfaces such as web browsers (e.g., web-based email). Aside from limited user-specific application configuration settings, consumers neither manage nor control the underlying cloud infrastructure, including the network, servers, operating system, storage, or even individual application capabilities.

[0099] Platform as a Service (PaaS): This provides consumers with the ability to deploy consumer-created or acquired applications on cloud infrastructure using programming languages ​​and tools supported by the provider. Consumers neither manage nor control the underlying cloud infrastructure, including networks, servers, operating systems, or storage, but they have control over the applications they deploy and may also have control over the configuration of the application hosting environment.

[0100] Infrastructure as a Service (IaaS): This provides consumers with the capability to deploy and run any software, including operating systems and applications, on the cloud, providing them with processing, storage, networking, and other basic computing resources. Consumers neither manage nor control the underlying cloud infrastructure, but they have control over the operating system, storage, and deployed applications, and may have limited control over chosen network components (e.g., host firewalls).

[0101] The deployment model is as follows:

[0102] Private cloud: A cloud infrastructure that runs exclusively for a single organization. It can be managed by that organization or a third party, and can exist inside or outside the organization.

[0103] Community cloud: A cloud infrastructure shared by several organizations and supporting a specific community with common interests (e.g., mission, security requirements, policies, and compliance considerations). It can be managed by the organization or a third party and can exist inside or outside the organization.

[0104] Public cloud: Cloud infrastructure available to the general public or large industrial groups and owned by organizations that sell cloud services.

[0105] Hybrid cloud: A cloud infrastructure consisting of two or more clouds (private, community, or public) that remain distinct entities but are bound together by standardized or proprietary technologies that enable data and applications to be ported together (e.g., cloud bursts for load balancing between clouds).

[0106] Cloud computing environments are service-oriented, characterized by statelessness, loose coupling, modularity, and semantic interoperability. The core of computing is the infrastructure comprising a network of interconnected nodes.

[0107] Now for reference Figure 14The diagram illustrates an illustrative cloud computing environment 1400. As shown, the cloud computing environment 1400 includes one or more cloud computing nodes 1402 to which local computing devices used by cloud consumers can communicate. These local computing devices include, for example, personal digital assistants (PDAs) or cellular phones 1404, desktop computers 1406, laptop computers 1408, and / or automotive computer systems 1410. The nodes 1402 can communicate with each other. They can be physically or virtually grouped (not shown) in one or more networks, such as private clouds, community clouds, public clouds, or hybrid clouds, or combinations thereof, as described above. This allows the cloud computing environment 1400 to provide Infrastructure as a Service, Platform as a Service, and / or Software as a Service, without requiring cloud consumers to maintain resources on their local computing devices. It should be understood that... Figure 14 The types of computing devices 1404-1410 shown are intended to be illustrative only. Computing node 1402 and cloud computing environment 1400 can communicate with any type of computerized device on any type of network and / or network-addressable connection (e.g., using a web browser).

[0108] Now for reference Figure 15 This demonstrates the 1400 (cloud computing environment) Figure 14 This provides a set of functional abstraction layers. For brevity, repeated descriptions of similar elements used in other embodiments described herein are omitted. It should be understood beforehand that... Figure 15 The components, layers, and functions shown are intended to be illustrative only, and embodiments of the invention are not limited thereto. As shown, the following layers and corresponding functions are provided.

[0109] The hardware and software layer 1502 includes hardware and software components. Examples of hardware components include: a mainframe 1504; a server 1506 based on a RISC (Reduced Instruction Set Computer) architecture; a server 1508; a blade server 1510; a storage device 1512; and a network and networking component 1514. In some embodiments, the software components include network application server software 1516 and database software 1518.

[0110] The virtualization layer 1520 provides an abstraction layer from which the following examples of virtual entities can be provided: virtual server 1522; virtual storage 1524; virtual network 1526, including virtual private network; virtual application and operating system 1528; and virtual client 1530.

[0111] In one example, management layer 1532 can provide the functions described below. Resource provisioning function 1534 provides dynamic acquisition of computing resources and other resources for performing tasks in the cloud computing environment. Metering and pricing function 1536 provides cost tracking for the use of resources in the cloud computing environment and provides bills or invoices for the consumption of these resources. In one example, these resources may include application software licenses. Security function provides authentication for cloud consumers and tasks and protection for data and other resources. User portal function 1538 provides access to the cloud computing environment for consumers and system administrators. Service level management function 1540 provides cloud resource allocation and management to meet required service levels. Service level agreement (SLA) planning and enforcement function 1542 provides pre-scheduling and procurement of cloud resources according to the SLA for its projected future needs.

[0112] Workload layer 1544 provides examples of functionalities that can leverage a cloud computing environment. Examples of workloads and functionalities that can be provided in this layer include: map creation and navigation 1546; software development and lifecycle management 1548; virtual classroom instruction provision 1550; data analysis and processing 1552; transaction processing 1554; and root cause analysis 1556. Various embodiments of the invention can be found using references. Figure 14 and 15 The cloud computing environment described is used to collect time-series data about one or more mechanical systems and / or perform one or more root cause analyses based on the time-series data.

[0113] This invention can be a system, method, and / or computer program product at any possible level of technical detail integration. A computer program product may include one or more computer-readable storage media having computer-readable program instructions thereon for causing a processor to execute aspects of the invention. A computer-readable storage medium may be a tangible device capable of holding and storing instructions for use by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable optical disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices such as punch cards or recessed structures with instructions recorded thereon, and any suitable combination of the foregoing. As used herein, computer-readable storage media should not be construed as temporary signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0114] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a suitable computing / processing device, or downloaded via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network) to an external computer or external storage device. The network may include copper cables, optical fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to a computer-readable storage medium within the respective computing / processing device.

[0115] Computer-readable program instructions used to perform the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, integrated circuit configuration data, or source code or object code written in any combination of one or more programming languages ​​(including object-oriented programming languages ​​such as Smalltalk, C++, etc.) and procedural programming languages ​​(such as the "C" programming language or similar programming languages). The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer via any type of network (including a local area network (LAN) or a wide area network (WAN)) or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuits including, for example, programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs) may execute computer-readable program instructions by utilizing state information from the computer-readable program instructions to personalize the electronic circuits in order to perform aspects of this invention.

[0116] Various aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0117] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / actions specified in one or more blocks of a flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other device to operate in a particular manner, such that the computer-readable storage medium in which the instructions are stored includes an article of writing comprising instructions for implementing aspects of the functions / actions specified in one or more blocks of a flowchart and / or block diagram.

[0118] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer-implemented process, such that the instructions executed on the computer, other programmable apparatus or other device perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0119] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions comprising one or more executable instructions for implementing a specified logical function. In some alternative embodiments, the functions indicated in the blocks may occur in a non-consecutive order as shown in the figures. For example, two blocks shown consecutively may actually be executed substantially simultaneously, or these blocks may sometimes be executed in reverse order, depending on the functions involved. It will also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified function or action or executes a combination of dedicated hardware and computer instructions.

[0120] To provide additional context for the various embodiments described herein, Figure 16 The following discussion is intended to provide a general description of a suitable computing environment 1600 in which various embodiments of the embodiments described herein may be implemented. Although the embodiments have been described above in the general context of computer-executable instructions that can run on one or more computers, those skilled in the art will recognize that the embodiments may also be implemented in combination with other program modules and / or as a combination of hardware and software.

[0121] Typically, program modules include routines, programs, components, data structures, etc., that perform specific tasks or implement specific abstract data types. Furthermore, those skilled in the art will understand that the methods of this invention can be implemented using other computer system configurations, including single-processor or multi-processor computer systems, minicomputers, mainframes, Internet of Things (IoT) devices, distributed mechanical systems, and personal computers, handheld computing devices, microprocessor-based or programmable consumer electronics, etc., each of which can be operatively coupled to one or more associated devices.

[0122] The embodiments illustrated herein can also be practiced in a distributed computing environment, where certain tasks are performed by remote processing devices linked via a communication network. In a distributed computing environment, program modules can reside in both local and remote memory storage devices. For example, in one or more embodiments, a computer-executable component can be executed from memory that may include or consist of one or more distributed memory cells. As used herein, the terms “memory” and “memory cell” are used interchangeably. Further, one or more embodiments described herein can execute the code of a computer-executable component in a distributed manner, for example, by multiple processors working together or cooperating to execute code from one or more distributed memory cells. As used herein, the term “memory” can encompass a single memory or memory cell at one location or multiple memories or memory cells at one or more locations.

[0123] Computing devices typically include a variety of media, which may include computer-readable storage media, machine-readable storage media, and / or communication media, these two terms being used differently from each other herein. A computer-readable storage medium or a machine-readable storage medium can be any available storage medium accessible by a computer, and includes volatile and non-volatile media, removable and non-removable media. By way of example and not limitation, a computer-readable storage medium or a machine-readable storage medium may be implemented in conjunction with any method or technique used for storing information such as computer-readable or machine-readable instructions, program modules, structured data, or unstructured data.

[0124] Computer-readable storage media may include, but is not limited to, random access memory (“RAM”), read-only memory (“ROM”), electrically erasable programmable read-only memory (“EEPROM”), flash memory or other memory technologies, optical disc read-only memory (“CD-ROM”), digital versatile disc (“DVD”), Blu-ray disc (“BD”) or other optical disc storage, magnetic tape cassettes, magnetic tape, disk storage or other magnetic storage devices, solid-state drives or other solid-state storage devices, or other tangible and / or non-transitory media that can be used to store desired information. In this regard, the terms “tangible” or “non-transitory” as used herein for storage devices, memories, or computer-readable media shall be understood to exclude only the propagation of transient signals themselves as a modifier, and shall not waive the rights to all standard storage devices, memories, or computer-readable media that do not only propagate transient signals themselves.

[0125] Computer-readable storage media can be accessed by one or more local or remote computing devices, for example via access requests, queries or other data retrieval protocols, for various operations concerning the information stored on the media.

[0126] Communication media typically embody computer-readable instructions, data structures, program modules, or other structured or unstructured data in data signals (such as modulated data signals, e.g., carrier waves) or other transmission mechanisms, and include any information delivery or transmission medium. The term "modulated data signal" or multiple signals refers to a signal whose one or more characteristics are set or altered in a manner that encodes information in one or more signals. By way of example and not limitation, communication media include wired media (such as wired networks or direct-wire connections) and wireless media (such as acoustic, RF, infrared, and other wireless media).

[0127] Refer again Figure 16 An example environment 1600 for implementing embodiments of the aspects described herein includes a computer 1602, which includes a processing unit 1604, system memory 1606, and a system bus 1608. The system bus 1608 couples system components, including but not limited to system memory 1606, to the processing unit 1604. The processing unit 1604 can be any of a variety of commercially available processors. Dual-microprocessor and other multiprocessor architectures may also be used as the processing unit 1604.

[0128] System bus 1608 can be any of several types of bus architectures, which can further interconnect to memory buses (with or without memory controllers), peripheral buses, and local buses using any of a variety of commercially available bus architectures. System memory 1606 includes ROM 1610 and RAM 1612. The basic input / output system (“BIOS”) can be stored in non-volatile memory such as ROM, erasable programmable read-only memory (“EPROM”), or EEPROM, where the BIOS contains basic routines that help transfer information between components within computer 1602, such as during startup. RAM 1612 may also include high-speed RAM (such as static RAM) for caching data.

[0129] Computer 1602 further includes an internal hard disk drive (“HDD”) 1614 (e.g., EIDE, SATA), one or more external storage devices 1616 (e.g., floppy disk drive (“FDD”) 1616, memory stick or flash drive reader, memory card reader, etc.), and an optical disc drive 1620 (e.g., capable of reading from or writing to CD-ROMs, DVDs, BDs, etc.). Although the internal HDD 1614 is illustrated as being located within computer 1602, the internal HDD 1614 may also be configured for external use in a suitable rack (not shown). Additionally, although not shown in environment 1600, a solid-state drive (“SSD”) may be used in addition to or in place of the HDD 1614. The HDD 1614, one or more external storage devices 1616, and optical disc drive 1620 may be connected to system bus 1608 via HDD interface 1624, external storage interface 1626, and optical drive interface 1628, respectively. The interface 1624 for external driver implementation may include at least one or both of Universal Serial Bus (“USB”) and Institute of Electrical and Electronics Engineers (“IEEE”) 1394 interface technologies. Other external driver connectivity technologies are also within the scope of the embodiments described herein.

[0130] The drive and its associated computer-readable storage medium provide non-volatile storage of data, data structures, computer-executable instructions, etc. For computer 1602, the drive and storage medium accommodate any data stored in a suitable digital format. Although the above description of computer-readable storage media refers to a corresponding type of storage device, those skilled in the art will understand that other types of computer-readable storage media, whether currently existing or developed in the future, may also be used in the example operating environment, and further, any such storage medium may contain computer-executable instructions for performing the methods described herein.

[0131] Multiple program modules can be stored in the drive and RAM 1612, including an operating system 1630, one or more application programs 1632, other program modules 1634, and program data 1636. All or part of the operating system, application programs, modules, and / or data can also be cached in RAM 1612. The systems and methods described herein can be implemented using a variety of commercially available operating systems or combinations of operating systems.

[0132] Computer 1602 may optionally include emulation technology. For example, a hypervisor (not shown) or other intermediary may emulate the hardware environment of operating system 1630, and optionally, the emulated hardware may be compatible with... Figure 16The hardware shown differs. In this embodiment, the operating system 1630 may include one of a plurality of virtual machines (VMs) hosted at the computer 1602. Furthermore, the operating system 1630 may provide a runtime environment, such as the Java Runtime Environment or the .NET Framework, to the application 1632. A runtime environment is a consistent execution environment that allows the application 1632 to run on any operating system that includes that runtime environment. Similarly, the operating system 1630 may support containers, and the application 1632 may be in the form of a container, which is a lightweight, standalone executable package that includes, for example, code, runtime, system tools, system libraries, and application settings.

[0133] Furthermore, computer 1602 may enable security modules, such as a Trusted Processing Module (“TPM”). For example, with TPM, the boot component hashes to the next boot component in time and waits for the result to match a security value before loading the next boot component. This process can occur at any layer of the computer 1602’s code execution stack, for example, it can be applied at the application execution level or at the operating system (“OS”) kernel level, thereby achieving security at any code execution level.

[0134] Users can input commands and information to computer 1602 through one or more wired / wireless input devices (such as keyboard 1638, touchscreen 1640, and pointing devices such as mouse 1642). Other input devices (not shown) may include microphones, infrared (“IR”) remote controls, radio frequency (“RF”) remote controls, or other remote controls, joysticks, virtual reality controllers and / or virtual reality headsets, gamepads, pointers, image input devices (e.g., one or more cameras), gesture sensor input devices, visual motion sensor input devices, emotion or face detection devices, biometric input devices (e.g., fingerprint or iris scanners), etc. These and other input devices are typically connected to processing unit 1604 via input device interface 1644, which can be coupled to system bus 1608, but may also be connected via other interfaces (such as parallel ports, IEEE 1394 serial ports, game ports, USB ports, IR interfaces, etc.). (Interfaces, etc.) connections.

[0135] The monitor 1646 or other types of display devices may also be connected to the system bus 1608 via an interface (such as the video adapter 1648). In addition to the monitor 1646, the computer typically includes other peripheral output devices (not shown), such as speakers, printers, etc.

[0136] Computer 1602 can operate in a networked environment via a logical connection to one or more remote computers (such as remote computers 1650). Remote computers 1650 can be workstations, server computers, routers, personal computers, portable computers, microprocessor-based entertainment devices, peer-to-peer devices, or other common network nodes, and typically include many or all of the elements described relative to computer 1602, although only memory / storage device 1652 is shown for simplicity. The depicted logical connections include wired / wireless connections to a local area network (“LAN”) 1654 and / or a larger network (e.g., a wide area network (“WAN”)) 1656. Such LAN and WAN networking environments are common in offices and companies and facilitate enterprise-wide computer networks (such as intranets), all of which can connect to global communication networks such as the Internet.

[0137] When used in a LAN network environment, computer 1602 can connect to local area network 1654 via a wired and / or wireless communication network interface or adapter 1658. Adapter 1658 facilitates wired or wireless communication to LAN 1654, which may also include a wireless access point (“AP”) configured thereon for communicating with adapter 1658 in wireless mode.

[0138] When used in a WAN network environment, computer 1602 may include modem 1660, or may be connected to a communication server on WAN 1656 via other means (such as via the Internet) for establishing communication on WAN 1656. Modem 1660 may be built-in or external, and may be a wired or wireless device, and may be connected to system bus 1608 via input device interface 1644. In a networked environment, program modules described relative to computer 1602 or parts thereof may be stored in remote memory / storage device 1652. It is understood that the network connection shown is an example, and other means of establishing communication links between computers may be used.

[0139] When used in a LAN or WAN network environment, computer 1602 can access cloud storage systems or other network-based storage systems as a supplement to or replacement for external storage device 1616 as described above. Typically, the connection between computer 1602 and the cloud storage system can be established, for example, on LAN 1654 or WAN 1656 via adapter 1658 or modem 1660, respectively. When computer 1602 is connected to the associated cloud storage system, external storage interface 1626 can manage the storage provided by the cloud storage system with the help of adapter 1658 and / or modem 1660, just as it manages other types of external storage. For example, external storage interface 1626 can be configured to provide access to cloud storage sources as if these sources were physically connected to computer 1602.

[0140] Computer 1602 is operable to communicate with any wireless device or entity operatively configured for wireless communication (e.g., printer, scanner, desktop and / or portable computer, portable data assistant, communications satellite, any device or location associated with a wirelessly detectable tag (e.g., public telephone booth, newsstand, shelf, etc.), and telephone). This may include Wireless Fidelity (“Wi-Fi”) and Wireless technology. Therefore, communication can be a predefined structure like a regular network, or simply self-organizing communication between at least two devices.

[0141] The above description includes only examples of systems, computer program products, and computer-implemented methods. It is certainly impossible to describe every conceivable combination of components, products, and / or computer-implemented methods for the purpose of describing this disclosure; however, those skilled in the art will recognize that many further combinations and substitutions of this disclosure are possible. Furthermore, with regard to the use of the terms “comprising,” “having,” “possessing,” etc., in the detailed description, claims, appendices, and drawings, these terms are intended to be inclusive in a similar manner to how the term “comprising” is interpreted when used as a transitional word in the claims. Various embodiments have been described for illustrative purposes but are not intended to be exhaustive or limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope of the described embodiments. The terminology used herein has been chosen to best explain the principles of the embodiments, their practical application, or improvements to existing technologies in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A computer system, comprising: Memory, which stores computer-executable components; as well as A processor operatively coupled to the memory and executing the computer-executable components stored in the memory, wherein the computer-executable components include: The maintenance component employs a greedy hill-climbing process that uses a conditional independence tester to identify causal graphs on a set of variables from time-series data. The greedy hill-climbing process reduces the number of conditional independence tests performed by the conditional independence tester from an exponential number to a polynomial number to determine Granger causality between variables from time-series data of a mechanical system under a given set of conditions, thereby detecting the causes of failures in the mechanical system.

2. The system according to claim 1, further comprising: A partitioning component divides the time series data into multiple groups based on defined time intervals.

3. The system according to claim 2, further comprising: Structural components that generate multiple relational structures for the multiple groups based on association measurements defined by the conditional independence tester.

4. The system according to claim 3, further comprising: A matrix component that generates multiple adjacency matrices for the multiple relational structures, wherein the multiple adjacency matrices represent the relationships between the variables.

5. The system according to claim 4, further comprising: A clustering component that uses K-means clustering to cluster the plurality of adjacency matrices into two cluster groups.

6. The system according to claim 5, further comprising: An attack component identifies a first group from the plurality of groups, the first group being a member of a first cluster group from the two cluster groups and adjacent to a second group, the second group being a member of a second cluster group from the two cluster groups.

7. The system according to claim 6, further comprising: The cause component identifies variables where the Hamming distance between a first adjacency matrix and a second adjacency matrix is ​​greater than a defined threshold, wherein the first adjacency matrix is ​​derived from the plurality of adjacency matrices and represents the first group, and wherein the second adjacency matrix is ​​derived from the plurality of adjacency matrices and represents the second group.

8. A computer system, comprising: Memory, which stores computer-executable components; as well as A processor operatively coupled to the memory and executing the computer-executable components stored in the memory, wherein the computer-executable components include: The maintenance component detects the onset of failures in the mechanical system by employing a greedy hill-climbing process that uses a conditional independence tester to identify a causal graph on a set of variables from time-series data. The greedy hill-climbing process reduces the number of conditional independence tests of the conditional independence tester from an exponential number to a polynomial number to determine Granger causality between variables from time-series data of the mechanical system under a given set of conditions.

9. The system according to claim 8, further comprising: A partitioning component divides the time series data into multiple groups based on defined time intervals.

10. The system according to claim 9, further comprising: Structural components that generate multiple relational structures for the multiple groups based on association measurements defined by the conditional independence tester; as well as A matrix component that generates multiple adjacency matrices for the multiple relational structures, wherein the multiple adjacency matrices represent the relationships between the variables.

11. The system of claim 10, further comprising: A clustering component that uses K-means clustering to cluster the plurality of adjacency matrices into two cluster groups; as well as An attack component identifies a first group from the plurality of groups, the first group being a member of a first cluster group from the two cluster groups and adjacent to a second group, the second group being a member of a second cluster group from the two cluster groups.

12. A computer-implemented method, comprising: The system, operatively coupled to the processor, employs a greedy hill-climbing process that uses a conditional independence tester to identify causal graphs on a set of variables from time-series data. The greedy hill-climbing process reduces the number of conditional independence tests performed by the conditional independence tester from an exponential number to a polynomial number to determine Granger causality between variables from time-series data of a mechanical system under a given set of conditions, thereby detecting the causes of failures in the mechanical system.

13. The computer-implemented method according to claim 12, further comprising: The system divides the time series data into multiple groups based on defined time intervals.

14. The computer-implemented method according to claim 13, further comprising: The system generates multiple relational structures for the multiple groups based on the association measurements defined by the conditional independence tester; as well as The system generates multiple adjacency matrices for the multiple relational structures, wherein the multiple adjacency matrices represent the relationships between the variables.

15. The computer-implemented method according to claim 14, further comprising: The system uses K-means machine learning to cluster the multiple adjacency matrices into two cluster groups; The system identifies a first group from the plurality of groups, the first group being a member of a first cluster group from the two cluster groups and adjacent to a second group, the second group being a member of a second cluster group from the two cluster groups; as well as The system identifies variables whose Hamming distance between a first adjacency matrix and a second adjacency matrix is ​​greater than a defined threshold, wherein the first adjacency matrix is ​​derived from the plurality of adjacency matrices and represents the first group, and wherein the second adjacency matrix is ​​derived from the plurality of adjacency matrices and represents the second group.

16. A computer-implemented method, comprising: A system operatively coupled to a processor detects failures in a mechanical system by employing a greedy hill-climbing process that uses a conditional independence tester to identify a causal graph over a set of variables from time-series data. The greedy hill-climbing process reduces the number of conditional independence tests performed by the conditional independence tester from an exponential number to a polynomial number.

17. The computer-implemented method according to claim 16, further comprising: The system divides the time series data into multiple groups based on defined time intervals.

18. The computer-implemented method according to claim 17, further comprising: The system generates multiple relational structures for the multiple groups based on the association measurements defined by the conditional independence tester; as well as The system generates multiple adjacency matrices for the multiple relational structures, wherein the multiple adjacency matrices represent the relationships between the variables.

19. The computer-implemented method according to claim 18, further comprising: The system uses K-means clustering to cluster the multiple adjacency matrices into two cluster groups; as well as The system identifies a first group from the plurality of groups, the first group being a member of a first cluster group from the two cluster groups and adjacent to a second group, the second group being a member of a second cluster group from the two cluster groups.

20. A computer program product, the computer program product comprising: Computer program instructions that can be read by processing circuitry and executed by said processing circuitry to perform the method according to any one of claims 12 to 19.

21. A computer-readable storage medium storing computer program instructions that are loadable into the internal memory of a digital computer, including a software code portion that, when the program is run on the computer, performs the method according to any one of claims 12 to 19.

Citation Information

Patent Citations

  • Method and device for utilizing time sequence correlation to perform IT fault root cause analysis

    CN107301119A