Multi-model machine learning for root cause analysis using saliency maps

A two-model architecture using saliency maps enhances root cause analysis by accurately identifying industrial anomalies with reduced computational costs, improving operational reliability and decision-making.

JP2025160890APending Publication Date: 2025-10-23GENERAL ELECTRIC TECH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025058017
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-10
Filing Date
2025-03-31
Publication Date
2025-10-23

AI Technical Summary

Technical Problem

Existing machine learning systems struggle to accurately identify root causes of events, particularly in industrial processes, due to limitations in distinguishing causal deviations from symptomatic deviations, and often require high computational costs.

Method used

A two-model architecture is employed, where a first machine learning model predicts values of interest and generates prediction residuals, and a second model, trained on these residuals, generates a saliency map to pinpoint inputs contributing most to prediction errors, thereby identifying potential root causes.

Benefits of technology

This approach improves accuracy in root cause analysis and anomaly detection, reduces computational costs, and enhances operational reliability by enabling informed decision-making and reducing downtime in industrial processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025160890000001_ABST
    Figure 2025160890000001_ABST
Patent Text Reader

Abstract

To provide systems and methods, which enable functional improvement of computing techniques in comparison to alternate systems and methods for root cause analysis and anomaly detection.SOLUTION: Systems and methods are provided. A method includes providing, by a computing system comprising one or more computing devices, a plurality of input values to a first machine-learned model. The method further includes generating, by the computing system using the first machine-learned model based on the plurality of input values, a saliency map. In the method, the first machine-learned model is a model that was trained to predict a prediction residual associated with a second machine-learned model.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates generally to machine learning. More particularly, the present disclosure relates to systems and methods for performing machine-learned root cause analysis using two or more machine learning models. [Background technology]

[0002] Machine learning is a computer-implemented process that allows a computer to iteratively "learn" based on training data. For example, a computing system can be provided with training inputs, and the computing system can perform multiple operations on the training inputs to generate training outputs. The multiple operations can include parameterized operations, where the operations are based, at least in part, on adjustable parameters. The computing system can evaluate one or more training outputs and adjust one or more parameters based on the evaluation. This process can be repeated for multiple training iterations. The multiple operations or adjusted parameters "learned" during the training process can be referred to as a machine-learning model or a machine-learned model.

[0003] Root cause analysis is the analysis of one or more events to determine the root cause of the one or more events. For example, an event of interest (e.g., an industrial failure, a machine learning prediction error, an abnormal event, etc.) may have one or more causes, but in some cases, the cause of the event of interest may be the event itself, and the event itself may have its own cause. Such a chain of causes and events may be referred to as a causal chain. In some cases, the start or root of the causal chain may be referred to as a root cause. Summary of the Invention

[0004] Aspects and advantages of the systems and methods according to the present disclosure will be set forth in part in the description that follows, or will be obvious from the description, or may be learned by practice of the techniques.

[0005] According to one embodiment, a method is provided, including providing, by a computing system including one or more computing devices, a plurality of input values ​​to a second machine learning model. The method also includes generating, by the computing system, a saliency map based on the plurality of input values ​​using the second machine learning model, where the second machine learning model is a model trained to predict a prediction residual associated with the first machine learning model.

[0006] According to another embodiment, a computing system is provided. The computing system includes one or more processors. The computing system includes one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computer system to perform operations. The operations include providing a plurality of input values ​​to a second machine learning model. The operations include generating a saliency map based on the plurality of input values ​​using the second machine learning model. In the operations, the second machine learning model is a model trained to predict a prediction residual associated with the first machine learning model.

[0007] According to another embodiment, one or more non-transitory computer-readable media are provided. The non-transitory computer-readable media store instructions that, when executed by one or more processors, cause the one or more processors to perform operations, including providing a plurality of input values ​​to a second machine learning model. The operations include generating a saliency map based on the plurality of input values ​​using the second machine learning model. In the operations, the second machine learning model is a model trained to predict a prediction residual associated with the first machine learning model.

[0008] These and other features, aspects, and advantages of the present systems and methods will become better understood with reference to the following description and appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the present technology and, together with the description, serve to explain the principles of the technology.

[0009] A full and enabling disclosure of the present systems and methods, including the best mode of making and using the same, directed to one of ordinary skill in the art, is set forth in this specification, which makes reference to the accompanying figures. [Brief explanation of the drawings]

[0010] [Figure 1A] FIG. 1 is a block diagram of a first aspect of an exemplary system for training a machine learning model according to an embodiment of the present disclosure. [Figure 1B] FIG. 1 is a block diagram of a second aspect of an exemplary system for training a machine learning model according to an embodiment of the present disclosure. [Figure 2] FIG. 1 is a block diagram of an exemplary machine learning model according to an embodiment of the present disclosure. [Figure 3] FIG. 1 is a block diagram of an exemplary system for generating a saliency map according to an embodiment of the present disclosure. [Figure 4] FIG. 1 is a block diagram of an exemplary industrial application of root cause analysis, according to an embodiment of the present disclosure. [Figure 5] FIG. 1 is a diagram of exemplary time series data according to an embodiment of the present disclosure. [Figure 6] FIG. 1 is a flowchart diagram of an exemplary method according to an embodiment of the present disclosure. [Figure 7] FIG. 1 is a block diagram of an exemplary computing system according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0011] Reference will now be made in detail to the embodiments of the present systems and methods, one or more examples of which are illustrated in the drawings. Each example is provided by way of explanation of the present technology, not as a limitation thereof. Indeed, it will be apparent to those skilled in the art that modifications and variations can be made in the present technology without departing from the scope or spirit of the claimed technology. For example, features illustrated or described as part of one embodiment can be used in another embodiment to yield still a further embodiment. Accordingly, the present disclosure is intended to cover such modifications and variations as come within the scope of the appended claims and their equivalents.

[0012] The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any embodiment described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other embodiments. Additionally, unless otherwise specified, all embodiments described herein are to be considered exemplary.

[0013] In the detailed description, numerical and letter designations are used to refer to features in the drawings. Like or similar designations in the drawings and description are used to refer to like or similar parts of the invention. As used herein, the terms "first," "second," and "third" may be used interchangeably to distinguish one component from another, but do not denote the location or importance of the individual components.

[0014] Approximate terms such as "approximately," "about," "generally," and "substantially" are not intended to be limited to the exact value stated. In at least some cases, approximating language can correspond to the precision of an instrument for measuring a value or the precision of a method or machine for constructing or manufacturing a component and / or system. For example, approximating language can refer to within a margin of 1, 2, 4, 5, 10, 15, or 20% of an individual value, a range of values, and / or any of the endpoints defining the range of values. When used in the context of angles or directions, such terms include a range of plus or minus 10 degrees of the stated angle or direction. For example, "substantially perpendicular" includes directions within 10 degrees of perpendicular in any direction, e.g., clockwise or counterclockwise.

[0015] Terms such as "coupled," "fixed," and "attached," unless expressly stated otherwise herein, refer to both direct coupling, fixing, or attachment, and indirect coupling, fixing, or attachment via one or more intermediate components or features. As used herein, the terms "comprises," "comprising," "includes," "including," "has," and "having," or any other variation thereof, are intended to cover non-exclusive inclusions. For example, a process, method, article, or apparatus that includes a list of features is not necessarily limited to only those features but may include other features not expressly listed or that are inherent to such process, method, article, or apparatus. Furthermore, unless expressly stated to the contrary, "or" refers to an inclusive disjunction, not an exclusive disjunction. For example, condition A or B is satisfied by any one of the following: A is true (or present) and B is false (or absent), A is false (or absent) and B is true (or present), and both A and B are true (or present).

[0016] Herein, throughout the specification and claims, range limitations are combinable and interchangeable, and unless the context and language dictate otherwise, such ranges are identified and include all subranges subsumed therein. For example, all ranges disclosed herein are inclusive of the endpoints, and the endpoints are independently combinable with each other.

[0017] As used herein, the term "line" may refer to a pipe, hose, tube or other fluid-carrying conduit.

[0018] overview FIELD OF THE DISCLOSURE The present disclosure relates generally to systems and methods for machine-learned root cause analysis. More particularly, the present disclosure relates to systems and methods for using two or more machine learning models to perform root cause analysis using a saliency map. A saliency map may be, for example, a data structure that indicates which inputs to a machine learning model have the greatest impact on the machine learning model's output (i.e., which inputs are most "salient").

[0019] In some embodiments, a first machine learning model can be trained to predict a value of interest (e.g., industrial process data, sensor data, etc.). The trained first model can be used to generate training data for a second machine learning model. More specifically, the trained first model can generate a plurality of predictions, which can be compared (e.g., subtracted, etc.) with a plurality of ground truth values ​​to generate a plurality of prediction residuals (e.g., prediction errors). The second machine learning model can be trained to predict an expected prediction residual of the first machine learning model using the plurality of prediction residuals.

[0020] The second machine learning model can be used to perform root cause analysis. For example, inputs can be provided to the second machine learning model, and a saliency map can be generated. The saliency map can indicate which inputs to the second machine learning model most contribute to an increase in the expected prediction residual associated with the first machine learning model. These salient inputs can be considered to be potential root causes of the high prediction error of the first machine learning model or to be potential root causes of conditions associated with the high prediction error.

[0021] In some exemplary applications, such a two-model architecture can be used for root cause analysis of industrial faults. For example, a first machine learning model can be trained using data associated with the normal operation of an industrial process. A second machine learning model can be trained using data associated with both the normal operation of the industrial process and the abnormal (e.g., fault) operation of the industrial process. In some examples, a high absolute value prediction residual of a first machine learning model trained on normal operation data can be associated with an operational abnormality (e.g., an industrial fault). In such cases, a saliency map of the second machine learning model can be used to identify one or more root causes of the anomaly.

[0022] In some embodiments, the input data and training data can include time-series data. In some embodiments, the input data and training data can include data having multiple input channels. For example, the input data and training data can include time-series data associated with multiple timestamps, each timestamp being associated with multiple input values ​​for multiple input channels (e.g., associated with multiple industrial sensors, measurements, etc.). In some examples, the time-series data can be input to the second machine learning model as a sliding time window having a width of t timestamps, with m input channels per timestamp.

[0023] In some examples, the second machine learning model may have an architecture (e.g., a temporal convolutional network, etc.) configured to enable separation or extraction of channel-wise saliency associated with each input channel (e.g., its contribution to the expected absolute increase in the prediction residual of the first machine learning model). In some examples, the architecture of the second model may include one or more channel-wise layers (e.g., convolutional layers, etc.) and one or more inter-channel layers (e.g., convolutional layers, fully connected layers, etc.). For example, the channel-wise layer may include multiple operations, each of which receives as input multiple input values ​​from a single input channel and generates one or more outputs based solely on the input from that input channel. In this manner, for example, channel-wise saliency may be preserved through one or more layers, and channel-wise explainability may be improved compared to alternative model architectures.

[0024] The saliency map (e.g., per-channel saliency map) may be generated in any suitable manner (e.g., using a gradient-based method, a deconvolution method, etc.). In some examples, the saliency map may be generated by reversing one or more operations associated with one or more inter-channel layers. For example, the second machine learning model may process multiple input values ​​using per-channel layers and inter-channel layers to generate an embedding. Each layer may include, for example, multiple weights and one or more nonlinear activation functions (e.g., ReLU, etc.). Generating the saliency map may further include processing the embedding based on a transposed weight matrix including the weights of the inter-channel layer, for example, such that the weight operations of the inter-channel layer are reversed.

[0025] In some examples, the activation function may be configured to output zero for some input values, such that the embedding may include multiple zero-valued outputs of the inter-channel layer and multiple non-zero-valued outputs of the inter-channel layer. In this way, for example, per-channel saliency can be determined by converting non-zero-valued outputs of the inter-channel layer into per-channel contributions to the non-zero-valued outputs while ignoring contributions to inter-channel nodes with zero-valued outputs. However, zero-valued outputs are not required. For example, multiple smaller and larger outputs of the inter-channel layer can be converted proportionally into per-channel contributions to the smaller and larger outputs by inverting the inter-channel weighting.

[0026] In some examples, the architecture of the first model can be the same as or different from the architecture of the second model. For example, in some exemplary experiments according to the present disclosure, both the first and second machine learning models were temporal convolutional networks with the same architecture. However, any suitable model architecture can be used for the first machine learning model as long as the first machine learning model can generate suitable predictions of values ​​of interest (e.g., predictions of normal operating behavior of an industrial system).

[0027] In some example applications, the saliency map of the second machine learning model can be used to perform or recommend actions to identify, prevent, or correct root causes. For example, in an application associated with industrial failures, the computing system can identify root causes of past or predicted future industrial failures based on the saliency map and recommend maintenance actions (e.g., repairs, inspections, etc.) to correct or prevent the failures based on the identified root causes. In some examples, the computing system can automatically take actions to correct or prevent the failures. In other applications, the computing system can identify root causes of abnormal events or machine learning prediction errors based on the saliency map, determine corrective actions based on the root causes, and take the corrective actions.

[0028] Systems and methods according to exemplary aspects of the present disclosure may provide various technical effects and benefits. For example, in some examples, the provided systems and methods may improve accuracy in identifying root causes compared to alternative systems and methods. As another example, in some examples, the provided systems and methods may improve accuracy in detecting anomalies in industrial processes compared to alternative systems and methods. In some examples, the provided systems and methods may provide similar accuracy at reduced computational costs compared to alternative systems and methods. Furthermore, the provided machine learning architectures (e.g., provided temporal convolutional networks) may, in some examples, provide additional advantages compared to some alternative model architectures, such as advantages in parallel processing, adaptability, scalability, and mitigation of vanishing gradient problems. In this manner, for example, the exemplary systems and methods according to aspects of the present disclosure may improve the functionality of the computing system itself. Furthermore, improved accuracy in fault detection and root cause detection may, in some examples, improve the operational reliability of industrial processes, minimize downtime, prevent damage to industrial components, reduce maintenance costs, or provide data for informed decision-making regarding future industrial processes across various industrial sectors.

[0029] In exemplary experiments according to the present disclosure, the provided system and method were compared with alternative systems and methods for root cause analysis and anomaly detection. In the experiments, the systems and methods according to exemplary aspects of the present disclosure provided improved accuracy compared to the alternative methods. For example, in experiments in which each tested system ranked multiple input channels from 1 (most likely to be the root cause) to 51 (least likely to be the root cause), the provided system and method achieved an average true root cause rank of 1.99, compared to 8.59 for the best-performing alternative tested and 15.98 for the less-performing alternative tested. In some instances, the accuracy advantage of the provided system and method was particularly strong in experiments in which small input deviations associated with the true root cause resulted in large symptomatic input deviations in downstream channels. Thus, the provided system and method can overcome the shortcomings of traditional single-model or deviation-based approaches, which may, in some instances, fail to adequately distinguish causal deviations from symptomatic deviations. In additional exemplary experiments involving an anomaly detection task, the provided systems and methods achieved an area under the precision-recall curve of 0.9315, compared to 0.9305 for the best-performing alternative embodiment tested and 0.9240 for the worst-performing alternative embodiment tested.

[0030] Furthermore, it will be understood that in some examples, the performance of a machine learning model may be correlated with the computational cost associated with the model. For example, in some examples, increasing the size (e.g., number of parameters, etc.) or complexity of a machine learning model can improve performance, but also increase the computational cost (e.g., training cost, inference cost, etc.) of the model. Similarly, reducing the size of a machine learning model can reduce the computational cost (e.g., electricity cost, memory cost, etc.) associated with the machine learning model, but also reduce the model's performance (e.g., root cause identification accuracy, etc.). As another example, reducing the size (e.g., number of training iterations, etc.) of a training dataset can reduce the computational cost (e.g., electricity cost, memory cost, etc.) associated with training a machine learning model, but also reduce the model's performance (e.g., root cause identification accuracy, etc.). Thus, it will be understood that systems and methods that can provide higher accuracy at similar computational cost compared to alternative methods can also be configured to provide similar accuracy at reduced computational cost compared to alternative methods (e.g., by reducing the size of the model, etc.). In this way, for example, the provided systems and methods can improve the functionality of computing technology by providing similar technical performance at reduced computational cost compared to alternative methods.

[0031] Exemplary System 1A and 1B show block diagrams of two exemplary systems for training machine learning models according to exemplary aspects of the present disclosure. FIG. 1A shows a first machine learning model 108 being trained, and FIG. 1B shows a second machine learning model 122 being trained based on the prediction residuals of the first machine learning model 108. In some embodiments, the trained second machine learning model 122 can be used for root cause analysis, as described further below with respect to other figures.

[0032] 1A shows a first machine learning model 108 being trained. A training system 104 can provide input 106 to the first machine learning model 108 based on a dataset that includes normal (e.g., non-anomalous) training data 102. Based on the input 106, the first machine learning model 108 can generate output 110. Based on the output 110 and the normal training data 102, the training system 104 can provide model updates 112 to train the first machine learning model 108.

[0033] The normal training data 102 generally can include or represent various types of data (e.g., numeric, binary, sensor data, audio, visual, text, etc.). The normal training data 102 can include one or many different types of data. In some examples, the normal training data 102 can include time-series data (e.g., including data from multiple timestamps). In some examples, the normal training data 102 can include multi-channel data having multiple input channels. In some examples, the input channels can include measurement channels (e.g., associated with metrics, sensors, gauges, industrial measurements, etc.). In some examples, the input channels can include control channels (e.g., associated with control valves, actuators, computerized controllers, etc.). In some examples, the normal training data 102 can include data associated with expected or non-anomalous behavior (e.g., of one or more systems, etc.). For example, the normal training data 102 can include data associated with non-anomalous behavior of an industrial process (e.g., normal operating behavior, etc.), an industrial system, a business process or system, a human process or system, a machine-learned process or system, a natural physical process or system, etc. In some examples, normal training data 102 may include other non-anomalous data (e.g., data associated with non-anomalous language examples, etc.).

[0034] Training system 104 may be or include one or more software, firmware, or hardware components configured to process normal training data 102, output 110, mixed training data 114, output 118, and output 124, and generate model updates 112 and 126. In some examples, training system 104 may be or include one or more computing systems or computing devices, such as the computing systems (e.g., computing system 702, computing device 704, computing system 724) shown below with respect to FIG.

[0035] The input 106 can generally include or represent various types of data. In some examples, the input 106 can include the normal training data 102 or share one or more characteristics with the normal training data 102. For example, the input 106 can have any of the characteristics described above with respect to the normal training data 102. In some examples, the training system 104 can process training examples of the normal training data 102 and extract the input 106 having multiple input channels and an expected or ground truth output to compare with the output 110. In some examples, the input 106 can include time series data. In some examples, the time series data can include data associated with a fixed number of timestamps, such as time series data associated with a sliding time window. Exemplary embodiments of the input 106 data are further described below with reference to FIG. 5.

[0036] The first machine learning model 108 may include one or more machine learning models. The first machine learning model 108 may include various model architectures. In some examples, an exemplary model architecture of the first machine learning model 108 may include a sequence processing model architecture (e.g., a convolutional neural network, a recurrent neural network, a long short-term memory, a selective structured state space model, a transformer, etc.). For example, the first machine learning model 108 may be configured to receive an input sequence and generate an output prediction (e.g., a numeric prediction, a sequence prediction, etc.). For example, the first machine learning model 108 may be configured to generate an output that predicts a value of interest based on the input sequence. In some examples, the first machine learning model 108 may have the same or a different architecture as the second machine learning model 122. For example, in principle, any machine learning model that can provide predictions (e.g., accurate predictions or high-quality predictions, etc.) associated with normal training data 102 may be used.

[0037] The output 110 may be a value (e.g., a prediction) generated by the first machine learning model 108 based on the input 106. The output 110 may generally include or represent various types of data. In some examples, the output 110 may be, include, or be associated with one or more numeric data components to be compared to ground truth values ​​to determine a prediction residual (e.g., a prediction error, etc.). For example, the output 110 may include a numeric prediction, a plurality of numeric predictions, a class prediction including a numeric probability value assigned to each class, a binary class prediction configured to be compared to a numeric (e.g., floating-point) ground truth value, a text prediction configured to be compared to a ground truth value based on a numeric metric (e.g., Levenshtein edit distance, etc.), or any other prediction configured to be compared to a ground truth value to determine a numeric prediction residual. In some examples, the output 110 may be associated with time series data (e.g., associated with a particular timestamp or a particular time window in a time series, etc.). In some examples, the output 110 may share one or more characteristics with the normal training data 102 or components thereof. For example, the normal training data 102 may include one or more expected or ground truth outputs to be compared to the output 110. Such expected outputs may have a data type similar to (e.g., the same as) or different from the data type of the output 110. In some examples, the output 110 may have any of the characteristics described above with respect to the normal training data 102.

[0038] The model update 112 may include, for example, updates to one or more parameters of the first machine learning model 108. For example, the model update 112 may include updating one or more parameters of the first machine learning model 108 to optimize an objective value. Optimizing the objective may include minimizing the value of a loss function, such as the difference (e.g., absolute difference, squared difference, etc.) between the output 110 and an expected or ground truth output associated with the input 106 used to generate the output 110. Such a difference may be referred to as a prediction residual. In this manner, for example, the training system 104 may train the first machine learning model 108 to more accurately predict one or more expected outputs associated with the normal training data 102.

[0039] The model update 112 may include various training or learning techniques, such as, for example, backpropagation. For example, an evaluation signal may be backpropagated from an output (or another source of an objective function) through the first machine learning model 108 to update one or more parameters of the first machine learning model 108 (e.g., based on a gradient of the evaluation signal with respect to the parameter value). In some examples, a system including one or more machine learning models may be trained end-to-end. Gradient descent techniques may be used to iteratively update parameters over several training iterations. In some implementations, performing backpropagation of errors may include performing truncated backpropagation over time. In some examples, training the first machine learning model 108 may include implementing several generalization techniques (e.g., weight decay, dropout, etc.) to improve the generalization ability of the first machine learning model 108 being trained. Various objective functions, such as mean squared error, likelihood loss, cross-entropy loss, hinge loss, contrastive loss, or various other loss functions, may be used in the model update 112.

[0040] FIG. 1B shows a second machine learning model 122 trained at least in part based on the output 118 of the first machine learning model 108. The training system 104 can receive a data set including mixed (e.g., abnormal and normal) training data 114, provide an input 116 to the first machine learning model 108, and the first machine learning model 108 can generate an output 118 based on the input 116. The training system 104 can determine a prediction residual associated with the output 118 based on a comparison of the output 118 with one or more ground truth values of the mixed training data 114. Next, the training system 104 can provide an input 120 to the second machine learning model 122, and the second machine learning model 122 can generate an output 124 based on the input. Based on the output 124 and the prediction residual associated with the output 118, the training system 104 can provide a model update 126 and train the second machine learning model 122.

[0041] In many respects, the method for training the second machine learning model 122 can be similar (e.g., the same) to the method for training the first machine learning model 108 described above with respect to FIG. 1A. However, instead of being trained to predict an expected output or ground truth output associated with the normal training data 102, the second machine learning model 122 can be trained to predict a prediction residual (e.g., prediction error, prediction loss, etc.) associated with the output 118 of the first machine learning model 108. In some examples, the second machine learning model 122 can also be trained using data (e.g., mixed training data 114) different from the data (e.g., normal training data 102) used to train the first machine learning model 108. For example, in some examples, the first machine learning model 108 can be trained with only normal data, and the second machine learning model 122 can be trained with a mixture of abnormal and normal data.

[0042] The mixed training data 114 may include, for example, normal data 102 and other data. In some examples, the mixed training data 114 may include data having any of the characteristics (e.g., data type, time series, multiple channels, etc.) described above with respect to the normal data 102. In some examples, the mixed training data may include both anomalous data and non-anomalous data. For example, the mixed training data 114 may include normal training data 102 associated with expected or non-anomalous behavior (e.g., of one or more systems) along with data associated with corresponding anomalous behavior (e.g., of the same or similar systems). For example, the mixed training data 114 may include normal training data 102 associated with non-anomalous behavior of an industrial process (e.g., normal operating behavior), an industrial system, a business process or system, a machine-learned process or system, a natural physical process or system, a human process or system, etc., and additional data associated with anomalous behavior of a similar or same process or system. In some examples, the mixed training data 114 may include normal training data 102 with other non-anomalous data (e.g., data associated with non-anomalous language examples), and may further include associated anomalous data.

[0043] Input 116 may be input 106, include input 106, or share one or more characteristics with input 106. For example, input 116 may have any of the characteristics described above with respect to input 106 or normal training data 102. Additionally, input 116 may include inputs associated with anomalous data, such as input training examples associated with mixed training data 114.

[0044] Output 118 may be output 110, include output 110, or share one or more characteristics with output 110. For example, output 118 may have any of the characteristics described above with respect to output 110.

[0045] The input 120 can be the input 116, include the input 116, or share one or more characteristics with the input 116. For example, in each training iteration, the input 120 can be similar to (e.g., the same as) the input 116. For example, if the first machine learning model 108 and the second machine learning model 122 are configured to take the same number and types of inputs (e.g., the same number of input channels in a multi-channel input, the same number of timestamps in a sliding time window associated with a time series input, etc.), then in each training iteration, the input 120 can be (or can be the same as) the input 116. Furthermore, the input 120 can have any of the characteristics described above with respect to the input 106 or the normal training data 102.

[0046] The second machine learning model 122 may include one or more machine learning models. The second machine learning model 122 may include various model architectures. In some examples, an exemplary model architecture of the second machine learning model 122 may include a sequence processing model architecture (e.g., a convolutional neural network, a recurrent neural network, a long short-term memory, a selectively structured state space machine, a transformer, etc.). For example, the second machine learning model 122 may be configured to receive an input sequence and generate an output prediction (e.g., a numeric predicted value, etc.). For example, the second machine learning model 122 may be configured to generate an output that predicts a prediction residual of the first machine learning model 108 based on the input sequence. In some examples, the second machine learning model 122 may have the same or a different architecture as the first machine learning model 108. In some examples, the second machine learning model 122 may be or include a temporal convolutional network. Other architectures are possible. For example, in principle, any model architecture that maintains a degree of explainability (e.g., per-channel explainability) through one or more layers can be used without departing from the scope of the present disclosure. Exemplary details of an exemplary architecture of the second machine learning model 122 are further described below with respect to FIG. 2.

[0047] The output 124 may be, for example, an output configured to predict a prediction residual associated with the first machine learning model 108. For example, the prediction residual of the first machine learning model 108 may include one or more numerical values ​​(e.g., a numerical metric indicating a prediction error, etc.), and the output 124 may include one or more numerical values ​​having a similar (e.g., the same) format as the prediction residual. For example, if the output 110, 118 is a single-channel or single-prediction output, the output 124 may be a single numerical value (e.g., a floating-point value, etc.). If the output 110, 118 includes multiple values ​​(e.g., probabilities of multiple classes, etc.), the output 124 may be a single value (e.g., based on an aggregate metric indicating an overall prediction residual) or may include multiple values ​​(e.g., predicting a prediction residual for each of multiple classes).

[0048] The model update 126 may include, for example, an update to one or more parameters of the second machine learning model 122. For example, the model update 126 may include updating one or more parameters of the first machine learning model 108 to optimize the value of an objective, such as a loss function. In particular, optimizing the objective may include minimizing the value of a loss function that includes a difference between an output 124 and a prediction residual associated with a corresponding output 118. For example, training the second machine learning model 122 may include, at each iteration of multiple training iterations, selecting an input 116 from the mixed training data 114, providing the input 116 to the first machine learning model 108 and receiving the output 118 based on the input 116, determining a prediction residual associated with the output 118 based on the output 118 and a ground truth value associated with the mixed training data 114, providing the input 116 to the second machine learning model 122 and receiving the output 124 based on the input 116, and performing the model update 126 configured to reduce the loss function that includes the difference between the output 124 and the prediction residual. Determining the prediction residual may include, for example, subtracting the output 118 from a predicted value or ground truth value (e.g., included in the mixture training data 114) to determine the difference. In some examples, determining the prediction residual may include performing an operation (e.g., absolute value, squaring, etc.) to convert the difference to a non-negative value. In other respects, the model update 126 may be similar to (e.g., the same as) the model update 112. For example, the model update 126 may have any of the characteristics (e.g., backpropagation, loss function, generalization, etc.) described above with respect to the model update 112. In this manner, for example, the training system 104 may train the second machine learning model 122 to more accurately predict the prediction residual associated with the first machine learning model 108.

[0049] 2 is a block diagram illustrating an example model architecture of the second machine learning model 122. The second machine learning model 122 may receive a multi-channel input 220 and process the input 220 with one or more per-channel layers 230 to generate a plurality of per-channel layer outputs 232. The second machine learning model 122 may process the per-channel layer outputs 232 with one or more inter-channel layers 234 to generate a plurality of inter-channel layer outputs 236. The second machine learning model 122 may process the inter-channel layer outputs 236 with one or more fully connected layers 238 to generate a final output, which in some examples may correspond to an expected prediction residual 240 of the first machine learning model 108 (e.g., when trained according to FIGS. 1A-1B ).

[0050] Multi-channel input 220 can be input 120, include input 120, or share one or more characteristics with input 120. For example, multi-channel input 220 can have any of the characteristics described above with respect to input 120. In some examples, multi-channel input 220 can have multiple input channels, each associated with one or more (e.g., multiple) input values. A channel can include, for example, a logical grouping of two or more inputs. For example, in some examples, multi-channel input 220 can include time-series data including a plurality of t timestamps, each timestamp including a plurality of m measurements respectively associated with m measurement channels. As a non-limiting, illustrative example, an industrial process can be monitored by m sensors, m sensor groups, m data loggers, etc., and each measurement channel can consist of t values ​​(e.g., t measurements, t recorded data points, t aggregate values ​​each determined based on the multiple measurements, etc.) associated with a particular sensor, sensor group, etc. over t time steps. However, other types of channels are possible. In some examples, the group associated with each channel may correspond to one or more similarities, shared performance, or other relationships between the channel's inputs. In some examples, one or more channels of interest may be defined based on one or more explainability goals. As a non-limiting, illustrative example, the second machine learning model 122 may be configured to identify the timestamp at which an anomaly first occurred by treating each timestamp as a channel. As another non-limiting, illustrative example, if the explainability goal includes narrowing down the root cause to a condition detectable by a specific machine or a specific sensor, the channels may be grouped such that each channel is associated with multiple input values ​​from one machine or one sensor (e.g., t input values ​​across multiple t timestamps).

[0051] The per-channel layer 230 may include or correspond to, for example, multiple nodes, filters (e.g., convolution filters or kernels), or other operations, where each operation of the per-channel layer 230 may be configured to receive input from only one input channel and generate an output value based on the input from only one input channel. In some examples, the per-channel layer 230 may include one or more nodes, filters, or other operations for each input channel, such that the total number of outputs of the per-channel layer 230 may be an integer multiple of m (e.g., the same or different integer multiples for different per-channel layers 230, etc.). In some examples, the per-channel layer 230 may be or include a per-channel convolutional layer (e.g., a conv1d layer, etc.). Other types of per-channel layers (e.g., a per-channel self-attention layer, etc.) are also possible. In some examples, the per-channel convolutional layer may include one or more filters, such as a filter for each of the m input channels. In other examples, the one or more filters may have weights shared across channels, provided that each operation performed using the filter is a per-channel operation (e.g., uses input from only one channel). In some examples, the number of output values ​​d associated with each node or channel of the per-channel layer 230 may be less than the number of input values ​​(e.g., t) associated with the channel. In some examples, the nodes or filters of the per-channel layer 230 may include one or more weight data structures, such as a vector, matrix, or tensor containing multiple weights. In some examples, the nodes or filters of the per-channel layer 230 may include one or more activation functions, such as a nonlinear activation function (e.g., a sigmoid function, a rectified linear unit (ReLU) function, a Gaussian error linear unit, or a softmax). In some examples, an activation function (e.g., a ReLU, etc.) may have an output value of 0 for multiple (e.g., an infinite multiple) possible input values ​​and a non-zero output value for multiple (e.g., an infinite multiple) possible inputs.In some examples, processing the multi-channel input 220 using the per-channel layer 230 may include, for each node associated with a respective channel, multiplying a plurality of inputs associated with that channel by a weight data structure (e.g., matrix multiplication, etc.) and passing the resulting value through one or more activation functions. If the per-channel layer 230 includes a convolutional layer, processing the multi-channel input 220 may include, for each of one or more respective filters associated with a respective channel, convolving the filter with a plurality of subsets of inputs associated with the channel. Convolving the filter with the plurality of subsets may include, for each subset, multiplying the inputs associated with the subset by a weight data structure associated with the filter (e.g., matrix multiplication, etc.) and processing the resulting value through one or more activation functions. The matrix multiplication may include, for example, performing a plurality of element-wise multiplications to generate a plurality of element-wise products and summing the element-wise products to generate one or more matrix entries associated with the matrix product.

[0052] The per-channel layer output 232 may include, for example, intermediate values ​​generated or used by the second machine learning model 122 in generating a final prediction. The per-channel layer output 232 may include, for example, a plurality of respective intermediate values, each intermediate value associated with exactly one input channel and generated based on input from exactly one input channel. In some examples, the per-channel layer output 232 may include a numeric value (e.g., a floating-point value, an integer value, a quantized value), a binary value, or another suitable data type. In some examples, the per-channel layer output 232 may be, include, or be referred to as a machine learning embedding or a per-channel embedding. In some examples, the per-channel layer output 232 may be stored in or correspond to a vector, matrix, or tensor format.

[0053] The inter-channel layer 234 may include, for example, one or more nodes, filters, or other operations configured to receive input values ​​associated with two or more channels (e.g., all channels) and generate an output based on the inputs associated with the two or more channels. In some examples, the inter-channel layer 234 may be an inter-channel convolutional layer (e.g., a conv1d layer, etc.) or may include an inter-channel convolutional layer. Other types of inter-channel layers 234 are also possible (e.g., an attention layer, a fully connected layer, etc.). In some examples, the inter-channel convolutional layer may include one or more filters that can convolve across multiple subsets of the per-channel layer output 232. One or more subsets (e.g., all subsets) over which such a filter is convolved may include inputs from two or more distinct channels. In an exemplary embodiment, the multi-channel input 220 may have m channels containing one or more measurements for each of t timestamps. The per-channel layer 230 may apply t×1 1D convolutions to generate d-dimensional embeddings for each of the m input channels. The inter-channel layer 234 may include a conv1D network with a filter size of t×md. In some examples, a node or filter in the inter-channel layer 234 may include one or more weight data structures, such as a vector, matrix, or tensor containing multiple weights. In some examples, a node or filter in the inter-channel layer 234 may include one or more activation functions, such as a nonlinear activation function. In some examples, an activation function (e.g., ReLU, etc.) may have an output value of 0 for multiple (e.g., infinite multiple) possible input values ​​and a non-zero output value for multiple (e.g., infinite multiple) possible inputs. In some examples, processing the per-channel layer outputs 232 using the inter-channel layer 234 may include, for each node in the inter-channel layer 234, multiplying the input associated with that node by a weight data structure (e.g., matrix multiplication, etc.) and passing the resulting value through one or more activation functions.In some examples (e.g., when the inter-channel layer 234 includes a convolutional layer), processing the multi-channel input 220 may include, for each of one or more filters of the inter-channel layer 234, convolving the filter across multiple subsets of the per-channel layer outputs 232. Convolving the filter across the multiple subsets may include, for each subset, multiplying the input associated with the subset by a weight data structure associated with the filter (e.g., matrix multiplication, etc.) and processing the resulting value through one or more activation functions.

[0054] The inter-channel layer output 236 may include, for example, intermediate values ​​generated by the inter-channel layer 234 based on the per-channel layer outputs 232. In some examples, the inter-channel layer output 236 may include numeric values ​​(e.g., floating-point values, integer values, quantized values), binary values, or other suitable data types. In some examples, the inter-channel layer output 236 may be, include, or be referred to as a machine learning embedding or an inter-channel embedding. In some examples, the inter-channel layer output 236 may be stored in or correspond to a vector, matrix, or tensor format.

[0055] The fully connected layer 238 may include, for example, a machine learning model layer (e.g., a neural network layer, etc.) including one or more nodes, each node of the fully connected layer connected to a respective input to the fully connected layer (e.g., each inter-channel layer output 236, etc.). For example, a node of the fully connected layer 238 may include multiple weights, and the number of weights of a node may be equal to the number of inter-channel layer outputs 236 received by the fully connected layer 238. In some examples, a node of the fully connected layer 238 may include one or more weight data structures, such as a vector, matrix, or tensor, including multiple weights. In some examples, a node of the fully connected layer 238 may include one or more activation functions, such as a nonlinear activation function. In some examples, processing the inter-channel layer outputs 236 using the fully connected layer 238 may include, for each node of the fully connected layer 238, multiplying the inter-channel layer output 236 by a weight data structure (e.g., matrix multiplication, etc.) and passing the resulting value through one or more activation functions.

[0056] The predicted first model residual 240 may be, for example, the final output of the second machine learning model 122, which in some examples may correspond to the expected prediction residual of the first machine learning model 108.

[0057] 3 is a block diagram illustrating an example system for generating a saliency map 360 associated with a second machine learning model 122. The second machine learning model 122 may have an architecture as shown in FIG. 2 , where one or more inter-channel layers 234 of the second machine learning model may include multiple inter-channel weights 350 and one or more inter-channel activation functions 352. An input (e.g., an abnormal input 320) may be processed in one or more per-channel layers 230 and inter-channel layers 234 of the second machine learning model 122 to generate an inter-channel layer output 236. The inter-channel layer output 236 may be processed with transposed inter-channel weights 354, which may reverse the process associated with the inter-channel weights 350, to generate per-channel saliency 356. The per-channel saliency 356 may be aggregated according to per-channel aggregation 358 to generate the saliency map 360.

[0058] The abnormal input 320 can be, include, or share one or more characteristics with the multi-channel input 220 or input 120. For example, the abnormal input 320 can have any of the characteristics described above with respect to the input 120 or the multi-channel input 220. In some examples, the abnormal input 320 can be an input 120, 220 associated with a condition of interest, such as a condition for which it is desired to find a root cause. For example, in some examples, the abnormal input 320 can be associated with an industrial malfunction, a machine learning prediction error of the first machine learning model 108, an unexpected event, or other event of interest for root cause analysis.

[0059] The inter-channel weights 350 may be or include, for example, inter-channel layer weights (e.g., weight matrices, etc.) For example, the inter-channel weights 350 may include weights such as those described above with respect to the inter-channel layer 234 of FIG.

[0060] The inter-channel activation function 352 may be or include, for example, an activation function of an inter-channel layer (e.g., ReLU, etc.) For example, the inter-channel activation function 352 may include an activation function such as those described above with respect to the inter-channel layer 234 of FIG.

[0061] The transposed inter-channel weights 354 may include, for example, a matrix transpose or other transformation of the inter-channel weights 350. In some examples, the transposed inter-channel weights 354 may be configured to reverse a process associated with the inter-channel weights 350. For example, if the inter-channel layer 234 is configured to perform a matrix multiplication between the per-channel layer outputs 232 and the inter-channel weights 350, the transposed inter-channel weights 354 may be configured to reverse the matrix multiplication. For example, generating the per-channel saliency 356 may include performing a first matrix multiplication on the per-channel layer outputs 232 with the inter-channel weights 350 to generate inter-channel weight values ​​351, processing the inter-channel weight values ​​351 with the inter-channel activation function 352 to generate the inter-channel layer outputs 236, and performing a second matrix multiplication on the inter-channel layer outputs 236 with the transposed inter-channel weights 354, the second matrix multiplication being the inverse of the first matrix multiplication.

[0062] The per-channel saliency 356 may include, for example, multiple numerical values ​​(e.g., floating-point values, integer or quantized values, etc.) or a group of numerical values ​​(e.g., a vector, a matrix, etc.). In some examples, the per-channel saliency 356 may indicate the contribution of one or more inter-channel layer outputs 236 or one or more per-channel layer outputs 232 to the magnitude of the predicted first model residual 240. In some examples, the per-channel saliency 356 may include one or more numerical saliency values ​​for each of the m input channels. In some examples, each input channel may be associated with multiple per-channel saliency 356. In some examples, the per-channel saliency 356 may have one or more dimensions that are the same as the corresponding dimensions of the per-channel layer outputs 232. For example, if the per-channel layer 230 outputs d dimensional embeddings for each of the m input channels, the corresponding set of per-channel saliency 356 may include d values ​​for each of the m input channels. In some examples, the process for generating the saliency map 360 may be repeated for multiple multi-channel inputs 220 or anomalous inputs 320, such as multiple (n-t+1) inputs 220, 320 associated with a sliding time window of width t associated with a time series of n timestamps, where n≧t. In such cases, the per-channel saliency 356 for each of the (n-t+1) time window positions may include, for example, d values ​​for each of the m input channels.

[0063] The per-channel aggregation 358 may include, for example, any suitable process for aggregating multiple per-channel saliency associated with a particular channel to generate a single saliency value for the channel. In some examples, the per-channel aggregation 358 may include determining, for each channel of the multiple input channels, one or more statistical aggregate values ​​(e.g., mean, median, mean absolute value, mean squared, l-norm, etc.) of the multiple per-channel saliency 356 associated with that channel. In some examples, the per-channel aggregation may include one or more additional actions, such as ranking or transforming the one or more per-channel saliency 356 or aggregate values.

[0064] The saliency map 360 may include, for example, multiple respective values ​​indicating the respective saliency of multiple respective channels. For example, the saliency map 360 may include, for each channel, a floating-point value (e.g., average channel-specific saliency 356, etc.), an integer value (e.g., a saliency rank associated with the channel, etc.), or other value indicating the respective saliency of the channel (e.g., at a particular time step or time window). In some examples, one or more highest-ranked or most salient input channels may be identified as the root cause of a state or event of interest. In some examples, the saliency map 360 may be combined with other saliency maps 360 associated with other time steps (e.g., associated with a sliding time window) to map the saliency of each of the multiple input channels over time. In some examples, an overall root cause may be determined based on multiple saliency maps 360 associated with multiple time steps (e.g., of a sliding time window). For example, in some examples, the d×m channel-specific saliency 356 for each window position of the sliding time window can be aggregated by calculating the average absolute contribution of d for each input channel to generate an m-channel saliency map 360 for each window position of the sliding time window. An overall rank across the entire time series can then be determined for each input channel by calculating the l2-norm of each input channel across all window positions of the sliding time window. In some examples, the input channel with the high (e.g., highest) l2-norm can be identified as the channel that is most likely (e.g., most likely to be) the root cause of the anomaly.

[0065] In some examples, one or more actions (e.g., maintenance actions such as repair or inspection actions) can be recommended or performed based on the salience map 360. For example, recommending or performing an action based on the salience map 360 can include determining a root cause based on the salience map and recommending or performing an action based on the determination. Recommending an action can include, for example, accessing a data structure (e.g., a database, a table, a file, etc.) that relates multiple root causes to multiple maintenance actions, retrieving data from the data structure associated with the root cause determined based on the salience map, determining a recommended action based on the retrieved data, and outputting the recommended action. Outputting the recommended action can include, for example, assigning the action to an entity (e.g., assigning to a device via an application programming interface, an electrical signal, a network signal, etc., assigning to a human via a workflow system, etc.), sending an action request, or outputting the recommended action (e.g., to a human user). Executing the action can include, for example, causing the action to be performed (e.g., by a computing device, a sensor device, an actuator, etc.). As an illustrative example, an action may include fully or partially opening or closing a valve to adjust flow rate, pressure, or other measured physical characteristic of an industrial process (e.g., a gas turbine process). Performing such an action may include, for example, sending a signal (e.g., an electrical signal, a network signal, etc.) to a valve actuator or other controller to cause the controller to open or close the valve. As another example, an action may include performing a test, and performing the action may include sending a signal to a device (e.g., a robotic device, etc.) equipped with one or more sensors to cause the device to perform the test.In some examples, recommending an action may include prompting a machine learning model (e.g., a language model; a multimodal model for processing language data, sensor data, audio or visual data, and / or other data types; an exploration augmentation generative model, etc.) with data indicative of the determined root cause (e.g., text, sensor data, etc.) and determining a recommended action using the machine learning model. In some examples, the recommended action may include one or more internal actions (e.g., processor actions, read / write actions to memory or storage, etc.) performed by the computing system or computing device that performed the root cause analysis.

[0066] FIG. 4 is a block diagram of an exemplary industrial application of root cause analysis according to embodiments of the present disclosure. In the exemplary industrial process, an industrial input 402 may be processed by one or more industrial components 404, 408, 410, 412, 414 in an industrial process flow 406 to generate one or more industrial outputs 416. Sensors 422 may measure one or more aspects of the industrial process to generate measurement data. In some embodiments, the measurement data may include multiple measurement channels (e.g., multiple sensors, types of sensors, etc.) that generate time-series data (e.g., each sensor 422 may take a measurement every few minutes, etc.). The measurement data may be used as an input (e.g., multi-channel input 220, anomaly input 320, etc.) to the second machine learning model 122. In some examples, a saliency map 360 may be generated and used to identify root causes (e.g., root causes of industrial operation failures, etc.). In some exemplary experiments according to the present disclosure, the exemplary root cause analysis method was tested using an exemplary chemical manufacturing process called the Tennessee Eastman Process.

[0067] The industrial inputs 402 may include any input that can be provided to an industrial process, including, but not limited to, materials (e.g., raw materials, natural resources, manufacturing materials, fuels, etc.), energy (e.g., electricity, heat, light, etc.), labor, or other inputs. In some exemplary experiments according to the present disclosure, the exemplary industrial process used for root cause analysis was the Tennessee Eastman Process, and the industrial inputs 402 included, among other things, input chemicals for chemical reactions.

[0068] The industrial components 404, 408, 410, 412, 414 may include, for example, process steps performed on the industrial input 402, or machines, tools, equipment, people, or other means for performing process steps on the industrial input 402. In some example experiments according to the present disclosure, the example industrial process used for root cause analysis was the Tennessee Eastman Process, where the first industrial component 404 was a reactor for performing a chemical reaction, the second industrial component 408 was a condenser, the third industrial component 410 was a separator, the fourth industrial component 412 was a compressor, and the fifth industrial component 414 was a stripper.

[0069] 4 depicts a particular number of industrial components connected in a particular order for purposes of illustration and explanation, the methods of the present disclosure are not limited to the particularly depicted order or arrangement. In principle, the methods of the present disclosure can be applied to any process (e.g., an industrial process) for determining root causes associated with the process (e.g., root causes of industrial failures, etc.). For example, the various industrial components 404, 408, 410, 412, 414 can be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.

[0070] The industrial process flow 406 may be, for example, a processing path followed by one or more industrial inputs 402. In some examples, the physical embodiment of the industrial process flow 406 may include lines (e.g., pipes, etc.) for transporting the industrial inputs 402; vehicles, conveyor belts, or other transportation devices; and fixed-place transitions that may perform two or more processing steps on the industrial inputs 402 without physically moving the industrial inputs 402 between processing steps. In some example experiments according to the present disclosure, the example industrial process used for root cause analysis was the Tennessee Eastman Process, in which at least some industrial process flow 406 components may include lines for transporting fluids.

[0071] The industrial output 416 may include, for example, any intended output or effect of an industrial process (e.g., product, energy such as electricity or heat, environmental changes such as cooling, or other outputs or effects). In some example experiments according to the present disclosure, the example industrial process used for root cause analysis is the Tennessee Eastman Process, and the industrial output 416 may include manufactured chemicals.

[0072] Purge 418 may include, for example, industrial process flow 406 for disposal of one or more items (e.g., by-products), dissipation of waste energy (e.g., heat, pressure), or other purging. Recycle feed 420 may include, for example, a line or other flow path for returning partially processed input material to a previous processing stage.

[0073] The sensor 422 may include, for example, any device or process for generating, recording, or storing data associated with the industrial process of FIG. 4 . For example, while the word “sensor” is used for illustrative and descriptive purposes, any system or method for generating input data can be used in principle. For example, in some examples, the sensor 422 data may include data associated with one or more inputs or control mechanisms (e.g., operator inputs for adjusting the industrial input 402). In some examples, the sensor 422 may include one or more computing devices for measuring, storing, or calculating data. As shown, the sensor 422 may generally be located throughout an area (e.g., a building, a container, a machine, a room, an area, etc.) in which the industrial process of FIG. 4 may occur, or the sensor 422 may be configured to monitor an area or activity throughout the area. In some exemplary experiments according to the present disclosure, the exemplary industrial process used for root cause analysis was the Tennessee Eastman Process, and the sensor 422 data included flow and feed rate data from various rate sensors, temperature and pressure data, chemical composition data, valve data associated with multiple respective control valves, and other related industrial data.

[0074] In some examples, the multi-channel input 220 may include a channel for each sensor 422 of the multiple sensors 422. For example, in an exemplary experiment associated with the Tennessee Eastman Process, the inputs 106, 116, 120, 220, 320 included 51 input channels generated by 51 sensors 422. In other examples, a single channel may include data from multiple sensors 422. In some exemplary experiments according to the present disclosure, the exemplary multi-channel input 220 included a sliding time window of t timestamps, with each channel including t data points from a single sensor 422. In some examples, a sliding time window having a width of t timestamps may be used to divide a time series of n timestamps into multiple sets of (n-t+1) multi-channel inputs 220, with each set of multi-channel inputs 220 associated with t consecutive timestamp subsets of the n timestamps.

[0075] In some examples, the systems and methods disclosed herein may be applied to one or more industrial processes associated with turbomachines or gas turbines. Turbomachines are utilized in various industries and applications for energy transfer. For example, a gas turbine engine generally includes a compressor section, a combustion section, a turbine section, and an exhaust section. The compressor section gradually increases the pressure of a working fluid entering the gas turbine engine and supplies the compressed working fluid to the combustion section. The compressed working fluid and fuel (e.g., natural gas) are mixed in the combustion section and combusted in a combustion chamber to generate high-pressure and high-temperature combustion gases. The combustion gases flow from the combustion section into the turbine section, where they expand to generate work. For example, the expansion of the combustion gases in the turbine section can rotate a rotor shaft connected to, for example, a generator, to generate electricity. The combustion gases then exit the gas turbine through the exhaust section.

[0076] In some exemplary embodiments, exemplary industrial components 404, 408, 410, 412, 414 may include gas turbines, sections or components of gas turbines (e.g., compression section, combustion section, turbine section, exhaust section, etc.), sub-components of gas turbines (e.g., turbine section components such as rotor blades, shafts, etc.), or industrial devices for use in combination with gas turbines (e.g., electrical transmission and power generation components) or components thereof, etc.

[0077] Example Input Data 5 is a diagram of exemplary time series data 506 according to an embodiment of the present disclosure. Each of multiple measurement channels can include multiple measurements at multiple timestamps. Each measurement channel can be plotted on a chart having a time axis 502 and a measurement axis 504 showing the measurement value of that channel for each indicated timestamp. In some examples, a sliding time window 508 can define the inputs (e.g., multi-channel input 220, anomalous input 320, etc.) provided to the second machine learning model 122.

[0078] The time axis 502 may be, for example, an axis for illustrating a domain of time. In some examples, the resolution of the time axis 502 may be discrete or continuous. For example, in some examples, the time axis 502 may include multiple timestamps associated with discrete intervals. For example, in some exemplary experiments according to the present disclosure, the sensor 422 collected industrial data every three minutes for 25 hours, resulting in 500 time steps at discrete three-minute intervals.

[0079] The measurement axis 504 may be, for example, an axis for indicating the magnitude of a measurement value associated with a particular measurement channel. A measurement or measurement channel may include a value measured by one or more sensors 422 (e.g., temperature, pressure, flow rate, chemical composition, volume or fluid level, etc.), a value collected by a data logger, a value input to one or more control devices (e.g., valves, computing devices, etc.), or any other relevant value.

[0080] Time series data 506 may be, for example, a series of data points (eg, measurements taken at particular times) that can be plotted on a time axis 502 and a measurement axis 504 .

[0081] The sliding time window 508 can be, for example, a window of t consecutive timestamps defining an interval on the time axis 502. The sliding time window 508 can, for example, define multiple subsets of the time series data 506, each of which can be used as the multi-channel input 220, the anomaly input 320, or the input 106, 116, 120. In some examples, the sliding time window 508 can have a constant width or a fixed width that includes a fixed number t of time steps (e.g., for use in a temporal convolutional network with a fixed input width, etc.). However, this is not required, and machine learning model architectures configured to process time windows 508 of variable width can be used without departing from the scope of the present disclosure.

[0082] Example results In some exemplary experiments according to the present disclosure, exemplary embodiments were tested using the Tennessee Eastman process dataset, which contains realistic simulation data for both fault-free and fault-prone operation of a chemical plant process. The experiments compared exemplary embodiments according to the present disclosure with alternative root cause analysis architectures, including 1D and 2D convolutional neural networks, deep autoencoders, transformer-based multivariate multistage prediction, and long short-term memory. In the exemplary experiments, the provided systems and methods outperformed the tested alternatives according to multiple performance metrics. For example, in an experiment in which each tested system ranked multiple input channels from 1 (most likely to be the root cause) to 51 (least likely to be the root cause), the provided systems and methods achieved an average true root cause rank of 1.99, compared to 8.59 for the best-performing alternative tested and 15.98 for the lower-performing alternative tested. In another exemplary experiment, the provided system and method achieved an area under the precision-recall curve of 0.9315 in an anomaly detection task, compared to 0.9305 for the best-performing alternative embodiment tested and 0.9240 for the worst-performing alternative embodiment tested.

[0083] Exemplary Methods 6 illustrates a flowchart of an exemplary method for root cause analysis, according to an exemplary embodiment of the present disclosure. While FIG. 6 illustrates steps performed in a particular order for purposes of illustration and explanation, the methods of the present disclosure are not limited to the particularly depicted order or arrangement. Various steps of the exemplary method 600 may be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.

[0084] At 602, the example method 600 may include, by a computing system comprising one or more computing devices, training a first machine learning model with a first training dataset. In some examples, the first machine learning model may be, include, or be included in the first machine learning model 108. In some examples, the first training dataset may be, include, or be included in the normal training data 102. In some examples, the example method 600 may include, at 602, using one or more systems or performing one or more actions described with respect to FIG. 1A .

[0085] At 604, the example method 600 may include providing, by a computing system, a plurality of respective inputs associated with a second training dataset to the first machine learning model. In some examples, the second training dataset may be, include, or be included in the mixed training data 114. In some examples, the example method 600 may include, at 604, using one or more systems or performing one or more actions described with respect to FIG. 1B .

[0086] At 606, the example method 600 may include generating, by the computing system, a plurality of respective predictions based on the plurality of respective inputs using the first machine learning model. In some examples, the respective predictions may be, include, or be included in the output 118. In some examples, the example method 600 may include, at 606, using one or more systems or performing one or more actions described with respect to FIG. 1B .

[0087] At 608, the example method 600 may include determining, by the computing system, a plurality of respective prediction residuals based on the plurality of respective predictions. In some examples, the example method 600 may include using one or more systems or performing one or more actions described with respect to FIG. 1B .

[0088] At 610, the example method 600 may include training, by a computing system, a second machine learning model to predict a prediction residual of the first machine learning model using the plurality of respective inputs and the plurality of respective prediction residuals. In some examples, the second machine learning model may be, include, or be included in the second machine learning model 122. In some examples, training the second machine learning model may include performing one or more model updates 126. In some examples, the example method 600 may include, at 610, using one or more systems or performing one or more actions described with respect to FIG. 1B .

[0089] At 612, the example method 600 may include providing, by a computing system, one or more inputs to the second machine learning model. In some examples, the inputs may be, include, or be included in the multi-channel input 220 or the anomalous input 320. In some examples, the example method 600 may include, at 612, using one or more systems or performing one or more actions described with respect to FIG.

[0090] At 614, the example method 600 may include generating, by the computing system, a saliency map of the second machine learning model based on the one or more inputs. In some examples, the saliency map may be, include, or be included in saliency map 360. In some examples, the example method 600 may include, at 614, using one or more systems or performing one or more actions described with respect to FIG.

[0091] Exemplary Computing Systems and Devices 7 is a block diagram of an exemplary computing system. Computing system 702 can include one or more computing devices 704, each of which can include a processor device 706, a memory device 712, a storage device 714, or an input / output device 716. Computing device 704 can include one or more machine learning models 718 (e.g., first machine learning model 108, second machine learning model 122), or portions thereof, which can be located, for example, in storage device 714 or memory device 712. Computing system 702 can be connected via network 720 to one or more other systems, such as one or more industrial systems 722 (e.g., as described with reference to FIG. 4), computing systems 724 (e.g., client computing systems, third-party computing systems, computing systems for controlling or monitoring industrial processes, etc.), or systems associated with one or more events of interest 726 for which root cause analysis is desired.

[0092] The computing system 702 may include any number of computing devices 704, such as one computing device 704 or many computing devices 704. For example, a computing system 702 for a small-scale training task or a computing system for performing inference with a machine learning model may use one or several computing devices 704. As another example, a computing system 702 for a large-scale training task (e.g., training a machine learning model with many parameters or training based on a large training dataset) may include many computing devices 704 performing parallel computing. Parallel computing may include decomposing a computational task into multiple subtasks (e.g., training iterations, subcomponents of a machine learning model, etc.) and assigning one or more respective subtasks to each of the multiple computing devices 704. Parallel computing may further include communicating results of the multiple subtasks among the computing devices 704 or aggregating the results of the subtasks to generate a final result of the computational task.

[0093] Computing device 704 may include any type of computing device, such as a server, workstation, desktop, laptop, virtual machine, mobile device, or other computing device.

[0094] The processor 706 may include one or more central processing units (CPUs) 708 and one or more application specific integrated circuits (ASICs) 710, such as, for example, an ASIC for performing floating-point operations (e.g., a GPU), an ASIC for performing matrix multiplication, an ASIC for performing machine learning or artificial intelligence tasks, or other ASICs. The CPU 708 may comprise, for example, any hardware configured to operate as a CPU (e.g., a microprocessor, a microcontroller, a soft-core processor, etc.).

[0095] The memory device 712 may include one or more non-transitory computer-readable storage media for, for example, temporarily storing data (e.g., to facilitate faster data access from the storage device 714). The temporary storage medium may include, for example, one or more memory devices such as high-bandwidth memory, random access memory (e.g., RAM, DRAM, SDRAM, DDR SRAM, etc.), virtual memory, cache memory, etc. The memory device 712 may include, for example, volatile memory, non-volatile memory, and semi-volatile memory.

[0096] The storage device 714 may include one or more non-transitory computer-readable storage media, for example, for persistent storage of data (e.g., including when the computing device 704 is powered off). A persistent storage device may include non-volatile storage such as read-only memory (e.g., ROM, PROM, EPROM, EEPROM, etc.), flash memory (e.g., NAND flash memory, etc.), magnetic storage (e.g., hard disk drive, floppy disk, etc.), optical storage (e.g., CD, DVD, Blu-Ray, etc.), or other non-volatile memory. In some examples, the storage device 714 may include volatile or semi-volatile memory coupled to a continuous power source (e.g., power grid power, backup battery, etc.) that provides the volatile memory with the ability to preserve data when the computing device 704 is shut down.

[0097] The input / output devices 716 may include, for example, any device or component for receiving input from or providing output to devices, systems, people, or other entities other than the computing device 704. The input / output devices 716 may include, for example, a network connection; a network card or network adapter for communicating over a network connection; human input / output devices such as a keyboard, mouse, display monitor, speakers, camera, microphone, or other input / output devices. In some examples, the input / output devices 716 may include one or more sensors, such as sensors 422 or other sensors associated with the industrial system 722 or the events of interest 726. In some examples, the input / output devices 716 may include one or more devices for communicating with such sensors.

[0098] The machine learning model 718 may include, for example, data that stores one or more parameters of the machine learning model; computer-readable instructions (e.g., source code, object code, etc.) that, when executed by one or more processors 706, cause the processor to perform one or more operations of the machine learning model; or any other data or components associated with the machine learning model (e.g., the first machine learning model 108, the second machine learning model 122, etc.).

[0099] Network 720 may be or include, for example, the Internet or any other network (e.g., a local area network, a wide area network, a peer-to-peer network, etc.) configured to transfer computer-readable data between computing devices. Network 720 may include, for example, wired connections, wireless connections, or a combination of both wired and wireless connections. In some examples, network 720 may be associated with one or more communication protocols for communicating over network 720, such as Transmission Control Protocol (TCP), Internet Protocol (IP), Hypertext Transfer Protocol (HTTP), User Datagram Protocol (UDP), Border Gateway Protocol (BGP), Address Resolution Protocol (ARP), etc. In some examples, network 720 may be associated with one or more security protocols for secure communication over network 720, such as Secure Sockets Layer (SSL) or Transport Layer Security (TLS) protocols.

[0100] The industrial system 722 may include, for example, a system associated with an industrial process for which fault monitoring or root cause analysis may be desired. For example, the industrial system 722 may include one or more systems such as those described above with respect to FIG. 4. The industrial system 722 may include, for example, multiple sensors 422 for collecting data (e.g., time series data) associated with the industrial process. Such data may be provided, for example, to one or more machine learning models 718 to perform root cause analysis associated with the industrial system 722 (e.g., in response to an industrial fault).

[0101] Computing system 724 may include any type of computing device, such as a workstation, a server, a laptop, a desktop, a mobile device, a virtual device, or other computing system. Computing system 724 may include, for example, a client computing system, a server computing system, or a third-party computing system. In some examples, computing system 724 may include any of the components or have any of the characteristics described above with respect to computing device 704. In some examples, computing system 724 may include a single computing device or multiple computing devices.

[0102] The events of interest 726 may include, for example, any event for which a root cause analysis may be desired. For example, events of interest may include: machine-learned prediction errors; unusual or unexpected events, such as events associated with abnormal measurements, outlier measurements, or measurements associated with high machine-learning prediction residuals; malfunctions or adverse events, such as industrial malfunctions, engineering or construction malfunctions (e.g., bridge collapses, etc.), power outages, natural disasters, man-made disasters, or any other events for which a root cause analysis may be desired.

[0103] 7 illustrates one exemplary configuration of a computing system that can be used to implement the present disclosure. Other computing system configurations can be used as well. For example, individual components shown can be omitted, rearranged, or added without departing from the scope of the present disclosure.

[0104] This written description uses examples to disclose the invention, including the best mode, and to enable any person skilled in the art to practice the invention, including making and using any device or system and performing any incorporated methods. The patentable scope of the invention is defined by the claims, and may include other examples that occur to those skilled in the art. Such other examples are intended to be within the scope of the claims if they include structural elements that do not differ from the literal language of the claims, or if they include equivalent structural elements that have no substantial differences from the literal language of the claims.

[0105] Further aspects of the invention are provided by the subject matter of the following clauses.

[0106] 1. A method of root cause analysis comprising: providing, by a computing system comprising one or more computing devices, a plurality of input values ​​to a second machine learning model; and generating, by the computing system, a saliency map based on the plurality of input values ​​using the second machine learning model, wherein the second machine learning model is trained to predict a prediction residual associated with the first machine learning model.

[0107] One or more of these methods where multiple inputs include time series data.

[0108] The method of one or more of these clauses, wherein the plurality of input values ​​includes measurements associated with a plurality of measurement channels.

[0109] The method of one or more of these clauses, wherein the second machine learning model includes at least one channel-specific layer.

[0110] The method of one or more of these clauses, wherein the channel-specific layer is a convolutional layer.

[0111] The method of one or more of these clauses, wherein the saliency map includes a plurality of per-channel saliencies indicating the contribution of each measurement channel to the prediction of the second machine learning model.

[0112] The method of one or more of these clauses, wherein the plurality of measurement channels includes a measurement channel associated with an industrial process.

[0113] The method of any one or more of these clauses, wherein the first machine learning model is trained to predict the outcome of an industrial process during normal operating behavior.

[0114] The method of one or more of these clauses, wherein the second machine learning model is trained using a training dataset including prediction residuals of the first machine learning model, the prediction residuals being determined based on data associated with both normal and abnormal operating behavior of the industrial process.

[0115] The method of one or more of these clauses, wherein the plurality of input values ​​includes one or more values ​​associated with abnormal behavior of the industrial process.

[0116] The method of one or more of these clauses, further comprising identifying one or more root causes associated with the abnormal operating behavior of the industrial process based on the saliency map.

[0117] The method of one or more of these clauses, further comprising determining, by a computing system, recommended maintenance actions associated with the one or more root causes based at least in part on the saliency map.

[0118] One or more of these clauses where the maintenance action involves repair or replacement.

[0119] One or more of these clauses how the maintenance action includes inspection.

[0120] The method of one or more of these clauses, wherein generating the saliency map includes processing, by a computing system, the input values ​​with at least one layer of a second machine learning model to generate a machine-learned embedding, the at least one layer including one or more weights and one or more activation functions; and processing, by the computing system, the embedding based at least in part on the one or more weights to generate the saliency map.

[0121] The method of one or more of these clauses, wherein the one or more weights comprise a weight matrix, and wherein processing based at least in part on the one or more weights comprises processing the embedding based on a transpose of the weight matrix.

[0122] The method of one or more of these clauses, wherein generating the saliency map further includes aggregating, by the computing system, a plurality of processed values, the processed values ​​being determined by processing the embeddings.

[0123] The method of one or more of these clauses, further including identifying, by the computing system, based on the saliency map, a cause associated with a high absolute value of an output of the second machine learning model, the output corresponding to a prediction residual associated with the first machine learning model.

[0124] A computing system including one or more processors and one or more non-transitory computer-readable media that collectively store instructions that, when executed by the one or more processors, cause the computer system to perform operations, the operations including performing one or more of the methods of these clauses.

[0125] One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations, the operations including performing one or more of the methods of these clauses. [Explanation of symbols]

[0126] 102 normal training data 104 Training System 106 Input 108 First Machine Learning Model 110 Output 112 model updates 114 Mixed Training Data 116 inputs 118 Output 120 inputs 122 Second Machine Learning Model 124 output 126 model updates 220 multi-channel input 230 Channel-specific demographics 232 Channel-specific layer output 234 Inter-channel layer 236 Inter-channel layer output 238 fully connected layer 240 Prediction residuals, first model residuals 320 Abnormal Input 350 Inter-channel weighting 351 Inter-channel weighting value 352 Inter-channel activation function 354 Transposed Inter-Channel Weights 356 Channel-specific salience 358 Channel Aggregation 360 Saliency Map 402 Industrial Input 404 First Industrial Component 406 Industrial Process Flow 408 Second Industrial Component 410 Third Industrial Component 412 Fourth Industrial Component 414 The fifth industrial component 416 Industrial Output 418 Purge 420 Recirculation Feed 422 Sensors 502 Timeline 504 Measurement Axis 506 Time Series Data 508 Sliding Time Window 600 Exemplary Method 702 Computing Systems 704 Computing Devices 706 Processor Device, Processor 708 CPU 710 ASIC 712 Memory Devices 714 Storage Devices 716 Input / Output Devices 718 Machine Learning Models 720 Network 722 Industrial Systems 724 Computing Systems 726 Interesting Events

Claims

1. providing (612) a plurality of input values ​​(120, 220, 320) to a first machine learning model (108) by a computing system (702) comprising one or more computing devices (704); generating (614), by the computing system (702), a saliency map (360) based on the plurality of input values ​​(120, 220, 320) using the first machine learning model (108); A method (600) for root cause analysis comprising: A method (600) for root cause analysis, wherein the first machine learning model (108) is trained (610) to predict a prediction residual associated with a second machine learning model (122).

2. The method (600) of claim 1, wherein the plurality of input values ​​(120, 220, 320) comprises time series data.

3. The method (600) of claim 1, wherein the plurality of input values ​​(120, 220, 320) comprises measurements associated with a plurality of measurement channels.

4. The method of claim 3 , wherein the first machine learning model comprises at least one channel-specific layer.

5. 5. The method of claim 4, wherein the channel-specific layer is a convolutional layer.

6. 4. The method of claim 3, wherein the saliency map comprises a plurality of per-channel saliencies indicating the contribution of each measurement channel to the predictions of the first machine learning model.

7. The method (600) of claim 3, wherein the plurality of measurement channels comprises measurement channels associated with an industrial process.

8. 8. The method of claim 7, wherein the second machine learning model is trained to predict outcomes of the industrial process during normal operating behavior.

9. 8. The method (600) of claim 7, wherein the first machine learning model (108) is trained (610) using a training dataset including prediction residuals of the second machine learning model (122), the prediction residuals being determined (608) based on data (114) associated with both normal and abnormal operating behavior of the industrial process.

10. The method (600) of claim 7, wherein the plurality of input values ​​(120, 220, 320) comprises one or more values ​​associated with abnormal behavior of the industrial process.

11. The method (600) of claim 7, further comprising identifying one or more root causes associated with abnormal operational behavior of the industrial process based on the saliency map (360).

12. 12. The method (600) of claim 11, further comprising determining, by the computing system (702), a recommended maintenance action associated with the one or more root causes based at least in part on the salience map (360).

13. The method (600) of claim 12, wherein the recommended maintenance action comprises repair or replacement.

14. The method (600) of claim 12, wherein the recommended maintenance action comprises an inspection.

15. Generating (614) the saliency map (360) comprises: processing, by the computing system (702), the input values ​​(120, 220, 320) with at least one layer (234) of the first machine learning model (108) to generate a machine-learned embedding, the at least one layer (234) including one or more weights (350) and one or more activation functions (352); processing, by the computing system (702), the embedding based at least in part on the one or more weights (350) to generate a saliency map (360); The method (600) of claim 1, comprising: