Providing robust root cause of analysis results output by AI-based analysis model

By receiving multivariate sensor data and generating robust meta-interpretive values ​​through weighted aggregation of interpretive models, the problem of machine learning models lacking global and local understanding in sensor data processing is solved, thereby improving the reliability and accuracy of anomaly detection and fault analysis.

CN121753015APending Publication Date: 2026-03-27SIEMENS AG
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-09
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing machine learning models lack a global and local understanding of the input data when processing sensor data, making it impossible to accurately determine the root cause of anomaly detection. Furthermore, inconsistencies exist between different interpretable AI methods when applied in parallel, affecting the reliability of anomaly detection and fault analysis.

Method used

By receiving multivariate sensor data and using a weighted aggregation of feature importance values ​​from at least two different interpretation models, robust meta-explanatory values ​​are generated, providing reliable root causes for the analysis results. Combining local and global interpretation models, meta-explanatory values ​​are generated to improve interpretability and accuracy.

Benefits of technology

It provides robust root causes for the output of machine learning models, improves the reliability and accuracy of anomaly detection and fault analysis, reduces processing power requirements, and generates high-quality interpretation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121753015A_ABST
    Figure CN121753015A_ABST
Patent Text Reader

Abstract

A computer-implemented method for providing a robust root cause of an analysis result (Yi) output by an AI-based analysis model (f) monitoring a machine (10) or a product of the machine (10), comprising:-receiving (S1) at least one collected data point (xi) comprising multivariable sensor data collected at the machine (10) or the product of the machine (10), each collected data point (xi) is configured as a multi-dimensional vector and each element of the vector represents a feature of the machine (10),-determining an analysis result (S2) with respect to the analysis task by processing the collected data points (xi) by an analysis model (f), -determining (S3) at least one meta-interpretation value indicative of at least one root cause of the determined analysis result (Yi) by weighted aggregation of feature importance values output by at least two different interpretation models (e1,..., en), where the weight value depends on the stability of the feature importance values at the collected data points (xi) for each of the interpretation models (e1,..., en), and-determining (S3) at least one meta-interpretation value indicative of at least one root cause of the determined analysis result (Yi) by weighted aggregation of the feature importance values output by the at least two different interpretation models (e1,..., en). And-outputting (S4) the analysis result (Yi) and a meta-interpretation value (mq) representing the root cause via the user interface.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to a computer-implemented method for providing robust root causes of analysis results output by an AI-based analysis model of input data of a processing machine or of products of a processing machine, a corresponding apparatus and a corresponding computer program product. BACKGROUND

[0002] Nowadays, sensors are ubiquitous in various kinds of heavy machinery, equipment of manufacturing plants and vehicles for monitoring and controlling autonomous driving. Data-driven applications using machine learning models process such sensor data and output an estimate of the state of the machine, equipment or observed environment.

[0003] One important field of application of sensors is the monitoring of the functionality of heavy machinery such as pumps, turbines, die casting machines, etc. For this purpose, sensors are installed on these equipment and machines and measure different physical parameters such as current, temperature, pressure, position and further this enables to monitor the state of the whole system. The sensor data sampled over time enables a continuous monitoring. If the machinery is subject to different damages, the sensor data values usually show unusual, suspicious patterns and anomalies in the data. This data can be analyzed by machine learning models which are trained to detect these anomalies.

[0004] Sensor data including image data is also used for quality monitoring of products manufactured by manufacturing equipment such as robots in manufacturing plants or additive manufacturing machines. Data collected by sensors and other equipment allowing to collect data related to the operation of one or several machines in a production line are input into all kinds of machine learning and AI techniques performing e.g. classification tasks and / or regression tasks, used for anomaly prediction and predictive maintenance of the machines or production line. Image data is processed in machine learning models performing image classification tasks for quality monitoring or object detection, controlling autonomous motion or driving. Further, regression tasks are applied for sensor fusion and prediction.

[0005] However, there are multiple problems associated with detecting anomalies and subsequent faults from sensor data. Since the state of a machine typically depends on various physical parameters and especially on combinations of specific parameters, a machine learning model has a complex structure since time series of sensor data representing all various parameters have to be considered to train the machine learning model. Moreover, such learned machine learning (ML) models are called black box models and the "logic" behind the provided results is typically opaque and not understandable by humans. This means that it is not possible to understand the global behavior of the model (i.e. how the model itself generally processes data) and it is not possible to locally understand how the machine learning model behaves, e.g. how the machine learning model makes a certain prediction based on specific input data. Throughout this document, a model is understood to be a machine learning model.

[0006] Not only are these local and global understandings needed to optimize a machine learning model, but also to provide more accurate outputs for various input data. The lack of global and local understanding of a machine learning model leads to problems in determining the influence of different measured physical parameters of the input data and subsequently inferring the root cause of a detected anomaly. In operational use, reliable information about the following features of the input is indispensable for taking countermeasures at the monitored machine or determining the root cause of an anomaly, which input features have a major contribution to the result output by the ML model.

[0007] There are different explainable Al methods which aim to make the reasoning of a machine learning model more understandable. Different explainable Al methods provide good results for different subsets of input data and can be applied to a ML model in parallel. These explainable Al methods are themselves based on different assumptions and constraints of the ML model. This makes them inherently conflicting and can be inconsistent. This is not only unfortunate but also destroys the value from applying different explainable techniques.

[0008] It is therefore an object of the present application to provide reliable and optimized information about root causes for results output by a ML model which resolves inconsistencies when multiple explainable Al methods are employed in parallel. SUMMARY

[0009] This object is solved by the features of the independent claims. The dependent claims contain further developments of the application.

[0010] A first aspect relates to a computer-implemented method for providing a robust root cause of an analysis result output by an AI-based analysis model monitoring a machine or a product of a machine, comprising the following steps: - receiving at least one collected data point comprising multivariate sensor data collected at a machine or a product of the machine, each collected data point being structured as a multi-dimensional vector and each element of the vector representing a feature of the machine, - determining an analysis result with respect to the analysis task by processing the collected data point by the analysis model, - determining a meta-explanation value indicating at least one root cause of the determined analysis result by a weighted aggregation of feature importance values output by at least two different explanation models, wherein the weight values depend on a stability of the feature importance values at the collected data point for each of the explanation models, and - outputting the analysis result and the meta-explanation value by a user interface.

[0011] The root causes of the determined analysis result comprise those features of the data point that have the highest importance values among all features of the data point and thus cause the determined analysis result. The meta-explanation value is one comprehensive importance value that takes into account all explanations provided by all applied explanation models for the considered data point but weighted with the stability of the explanation models at the data point. The meta-explanation value provides a reliable root cause for the analysis output of all data points and takes into account the explanations of all explanation models.

[0012] In an embodiment of the method, comprising the step of determining a feature importance value of each feature of the collected data point for the analysis result by each of the at least two different explanation models by inputting the ML analysis model and the collected data point into the explanation model.

[0013] Each of the explanation models analyzes the analysis model at the input data point and provides an importance value for each feature of the data point that indicates a contribution of this feature to the analysis result. This technique provides an explanation for the analysis result according to the importance of the contribution made by the features.

[0014] In an embodiment of the method, the weight values of the explanation models are determined by combining the importance values output by the explanation models when data points close to the collected data point are input.

[0015] Thus, the weight values provide not only the feature importance values at the input data point itself but also variations of the feature importance values output by the explanation models at additional data points located in the vicinity of the input data point of the analysis model. This provides a flexible qualitative assessment of the explanation models.

[0016] In an embodiment of the method, the weight values of the explanation models are determined by combining the importance values output by modified versions of the explanation models that differ by at least one hyperparameter of the explanation models.

[0017] Therefore, the weight values ​​take into account variations in the explanatory model itself. If several preferably slightly modified explanatory models provide similar feature importance values, the explanatory model's results have high confidence, and the weight values ​​are high. This provides an interpretive model with a meaningful criterion for providing reliable feature importance values.

[0018] In embodiments of the method, redundant explanatory models are identified by determining highly relevant explanatory models, and redundant explanatory models are applied for a longer period of time to determine feature importance values.

[0019] This reduces the processing power required to execute the method without significantly sacrificing the accuracy of the meta-interpreted values.

[0020] In an embodiment of the method, a global meta-interpretation value for all data points is generated by an aggregation method on all data points, preferably by averaging the weight values ​​on all data points.

[0021] This provides a measure to evaluate the quality of the results provided by a combination of explanatory modeling approaches. It also indicates whether the combination of explanatory models is suitable for the analytical model under consideration.

[0022] In embodiments of this method, the interpretation model is either a local interpretation model or a global interpretation model.

[0023] Therefore, evaluating meta-interpretations by combining local and global interpretation models provides a reliable and trustworthy underlying reason across a wide range of input data points.

[0024] In embodiments of this method, the local interpretation model and / or the global interpretation model are model-agnostic.

[0025] The model-agnostic interpretation means that the model is unaware of the specific analytical model. Therefore, this method can be applied to any analytical model.

[0026] In embodiments of this method, the analysis model is a machine learning model configured to perform classification or regression tasks, particularly a deep neural network or decision tree ensemble.

[0027] Combinatorial interpretive modeling approaches provide high-quality and reliable interpretations, especially for complex models such as deep neural networks or decision tree ensembles.

[0028] According to the present invention, alarm messages or instruction messages are derived from the analysis results and meta-interpretation values ​​and transmitted to the machine.

[0029] Alarm messages inform operators of the analysis results directly, especially if the results indicate anomalies, request short-term maintenance, or indicate low-quality products.

[0030] In embodiments of this method, the analysis task is one of anomaly detection, predictive maintenance prediction, or quality monitoring.

[0031] The second aspect relates to a monitoring device for analyzing a machine or a product of a machine, comprising: - A data interface configured to receive at least one collected data point, said data point including multivariable sensor data collected at the machine or the machine's product, each collected data point being constructed as a multidimensional vector, and each element of the vector representing a characteristic of the machine. - The analyzer unit is configured to determine the analysis results regarding the analysis task by processing the collected data points through the ML analysis model. - An interpreter unit configured to determine a meta-interpretive value indicating at least one root cause of the identified analytical outcome by weighted aggregation of feature importance values ​​output by at least two distinct interpretive models, wherein the weight values ​​depend on the stability of the feature importance values ​​at the collected data points for each of the interpretive models, and - User interface, which is configured to output analysis results and meta-explanation values ​​and / or root causes.

[0032] The monitoring device uses AI models to perform different tasks, thus providing not only task-related predictions but also the underlying causes, i.e., explanations of which features of the input data points primarily caused the predictions.

[0033] In addition, the device includes a machine interface configured to derive alarm messages and / or command messages from analysis results and meta-interpretation values, and to transmit alarm and / or command messages to the machine.

[0034] This device can directly control the monitored machine via command messages. Alarm messages can be analyzed by the machine and trigger further actions within it.

[0035] The third aspect relates to a computer program product that can be directly loaded into the internal memory of a digital computer, the computer program product including software code portions for performing the steps of the method for monitoring a machine as described above when the product is run on the digital computer.

[0036] Various types and kinds of machines can be considered. Examples include, for instance, factory equipment such as manufacturing apparatus, robots, or assembly line equipment. Further examples include power plants or transmission networks such as turbines or generators, and substation equipment. Examples include mobile infrastructure equipment such as public transportation vehicles, escalators, elevators, etc. Further examples include heavy machinery such as pumps, turbines, die-casting machines, etc. Further examples include vehicles such as automobiles, locomotives, airplanes, or ships. Attached Figure Description

[0037] The invention will be explained in more detail with reference to the accompanying drawings. Similar objects will be labeled with the same reference numerals.

[0038] Figure 1 An embodiment of the computer-implemented method of the present invention is illustrated by a flowchart.

[0039] Figure 2 An embodiment of the device of the present invention is illustrated schematically. Detailed Implementation

[0040] Note that in the following detailed description of the embodiments, the accompanying drawings are merely schematic, and the illustrated elements are not necessarily shown to scale. Rather, the drawings are intended to illustrate functions and the cooperation of functions. It should be understood here that any connection or coupling of functional blocks, devices, components, or other physical or functional elements may also be achieved through indirect connections or couplings, such as via one or more intermediate elements. Connections or couplings of elements or components or nodes may be achieved, for example, through wired connections, wireless connections, and / or a combination of wired and wireless connections. Functional units may be implemented by dedicated hardware (e.g., processors, firmware), or by software, and / or by a combination of dedicated hardware and firmware and software. It should also be noted that each functional unit described for the apparatus can perform functional steps of the associated method. Similarly, each functional step described for the method can be performed in a functional unit of the associated apparatus.

[0041] AI-based tools are applied to a wide variety of machines, equipment, and other types of technological systems to detect anomalies, predict maintenance needs, monitor the quality of manufactured products, or inspect objects. These AI-based tools include analytical models, which in most cases are black boxes that receive collected input data and output predictions. This black box provides no indication of which features of the input data have the greatest impact on the analytical model's output predictions and cause the model to produce the most accurate predictions. This leads to a problem of low user confidence in the predictions and responses (e.g., when anomalies are predicted).

[0042] The essence of interpretive methods lies in their aim to indicate which features of the input data are important for the analytical model's predictions, and in what way these features are important for the predictions. Throughout this document, interpretive methods are also referred to synonymously as interpretive models.

[0043] The proposed method provides a robust underlying reason for the analytical results output by an AI-based analytical model of a monitoring machine or its product. Details are provided by... Figure 1 Let me explain.

[0044] In the first step S1, at least one collected data point xi is received, which includes multivariate sensor data collected at the machine or its products. Each collected data point xi is constructed as a multidimensional vector, and each element of the vector represents a feature of the machine. Features are, for example, physical parameters collected by sensors, such as temperature, pressure, and power. Features of image data, such as sensor data from a vision sensor like a camera, are single pixels or subsets of pixels. Data points may also include sequences of sensor data collected over a specific time period. In this case, the number of features is given by multiplying the number of parameters in the input data points by the number of time points at which the sensor data was collected. A feature importance value is determined for each feature.

[0045] Regarding the analysis task, the analysis result Yi is determined by processing at least one collected data point xi by the analysis model f, see S2.

[0046] In step S3, a meta-explanatory value mq is determined. This meta-explanatory value mq indicates at least one root cause of the determined analysis result Yi by a weighted aggregation of feature importance values ​​output by at least two different explanatory models e1, ..., en, where the weight values ​​w depend on the stability of the feature importance values ​​at the collected data points xi for each of the explanatory models e1, ..., ee. The analysis result and the meta-explanatory value mq representing the root cause are output via the user interface, see step S4. The root cause of the analysis model's result indicates the most relevant feature of the input data that caused the analysis model to output the analysis result.

[0047] Feature importance is an interpretation method that identifies the most important features or variables in determining the output of a considered ML model (specifically, an analytical model). Feature importance is applied to the analytical model when processing specific input data points, and a feature importance value is assigned to each individual feature of the input data point. The assigned feature importance value provides a measure of the importance of the corresponding feature of the machine to the analytical result Yi output for a specific input data point xi.

[0048] To determine the meta-interpretive value mq in S3, the feature importance value of each feature of the collected data points xi to the analysis result Yi is determined by inputting the machine learning analysis model f and the collected data points xi into the interpretive models e1, ..., en, as shown in S31. This is performed for each of at least two different interpretive models e1, ..., en. The weight values ​​of the interpretive model ep are determined by combining the feature importance values ​​output by the interpretive models when inputting data points close to the collected data points, as shown in S32. Alternatively or additionally, the weight values ​​of each of the interpretive models are determined by combining the importance values ​​output by a modified version of the considered interpretive model, as shown in S33. The modified versions of the interpretive models differ by at least one hyperparameter of the interpretive model.

[0049] Redundant explanatory models are identified by determining highly correlated explanatory models. If the correlation indicator, which represents the degree of correlation between two explanatory models, is higher than a predefined threshold, one of the explanatory models is identified as redundant. Redundant explanatory models are no longer discarded and are no longer used to determine feature importance values.

[0050] Optionally, a global meta-interpretive value is generated by aggregating all data points. Preferably, the global meta-interpretive value is determined by averaging the weight values ​​across all data points. The global meta-interpretive value indicates the quality level of the combined interpretive model. In this combined interpretive approach, both local and global interpretive models can be applied. Preferably, each of the interpretive models is model-agnostic; that is, it provides an interpretation for any analytical model. Examples of local interpretive models are LIME (Simple Local Additive Approximation of Complex Models) and SHAP, which provide simple and additive interpretations for local instances. Local interpretive models provide interpretations of analytical results for a specific input data point, while global interpretive models provide interpretations of analytical results for a wide range of input data points.

[0051] Examples of global interpretation models are partial dependency graphs showing the partial influence of selected features on the output, or permutations of feature importance, SAGE, or impurity-based importance showing the partial importance of selected features.

[0052] The analytical model f is a machine learning model configured to perform a classification or regression task. Because the methods outlined here are model-agnostic, they are not constrained to any classifier or regression model. The analytical model can be a deep neural network (DNN) used for supervised methods, such as a DNN classifier like a Long Short-Term Memory (LSTM), a Convolutional Neural Network (CNN), or a Recurrent Neural Network (RNN) (if labels are available).

[0053] The analysis task is one of anomaly detection, predictive maintenance prediction, and quality monitoring. Alarm or instruction messages are derived from the analysis results and meta-interpreted values. Instruction messages are passed to the machine, preferably to the machine's controller, and applied to adapt to the machine's settings. Alarm messages are preferably passed to the machine's input interface, such as a monitoring panel indicating the most relevant characteristic that caused the provided analysis results. This could be an anomaly in the monitored machine's operation, a maintenance request, or product quality falling below a certain quality threshold.

[0054] exist Figure 2 An embodiment of apparatus 20 configured to perform the described method is depicted herein and is described below. Apparatus 20 is configured to analyze machine 10 or a product of machine 10. Apparatus 20 includes a data interface 21, an estimator unit 22, an interpretation unit 23, a user interface 24, and an optional machine interface 25.

[0055] Data interface 21 is configured to receive at least one collected data point xi, which includes multivariate sensor data collected at machine 10 or its products. Each collected data point xi is constructed as a multidimensional vector, and each element of the vector represents a feature of machine 10. Analyzer unit 22 is configured to determine an analysis result regarding the analysis task by processing the collected data points xi using a machine learning analysis model. Interpreter unit 23 is configured to determine a meta-explanation value indicating at least one root cause of the determined analysis result by weighted aggregation of feature importance values ​​output by at least two different interpretation models. User interface 24 is configured to output the analysis result and the meta-explanation value and / or root cause. Machine interface 24 is configured to derive a message M, i.e., an alarm message and / or instruction message, from the analysis result Yi and the meta-explanation value mq, and transmit the alarm and / or instruction message to machine 10.

[0056] The device 20 may preferably be one of an anomaly detector, a predictive maintenance assessor, or a quality monitor. The invention further includes a computer program product that can be directly loaded into the internal memory of a digital computer, comprising software code portions for performing the steps of the method described above when the product is run on the digital computer.

[0057] Device 20 monitors machine 10 based on input data collected from sensors at machine 10, and outputs information via user interface 24, namely, an analysis result Yi accompanied by an interpretation of that information (i.e., a meta-interpretation value mq). Users, such as monitoring personnel, use the analysis result and its interpretation to perform monitoring actions. In a more detailed embodiment, machine interface 25 of device 20 generates an alarm message M based on the analysis result Yi and the meta-interpretation value mq. This alarm message can be received and displayed by the monitoring equipment of machine 10.

[0058] In a more complex embodiment, machine interface 25 derives at least one instruction message from the analysis result Yi and the meta-interpretation value. The instruction message is passed to the machine and can cause changes to the settings of machine 10.

[0059] The following text provides a more detailed explanation of the basic considerations, central ideas, and specific methods.

[0060] Basic considerations Given input data and basic fact labels As an analytical model based on AI-based machine learning models Provide predictions, that is, output analysis results. Throughout this document and Figure 1 and Figure 2 Analysis output Also represented by Yi. Analysis model express The regression problem or and (in The classification problem involves predicting the regression task within a regression problem. Indicates the probability value of a specific event. In the context of solving classification problems, the analysis results... for Each of the possible categories is provided with a probability value between 0 and 1. Therefore, the disclosed method considers all standard supervised machine learning methods; that is, it considers all standard supervised analytics models. Analysis Model Not constrained to a specific use case. That is, the collected data points (It is constructed as a multidimensional vector, and each element of the vector represents a feature of the machine) and can represent tabular data, time-series data, or image data.

[0061] Explanation model To represent. Explanation model It can be one of the following two categories: • Attribution models based on local permutations, such as LIME (Locally Interpretable Models, Unpredictable Interpretations), SHAP (SHapley Additive Interpretations), and similar models.

[0062] • Measurements based on local gradients, such as ensemble gradients, Grad-Cam, saliency maps, and similar measurements.

[0063] In addition, globally interpretable AI models (i.e., global XAI methods) can also be used as interpretive models.

[0064] Explanation Model The key point is that they are designed to indicate which features, and in what way, are relevant to the analytical model. Prediction is important. If the collected data points are tabular data, then the features are the elements of the tabular data. If the collected data points are a time series of sensor data points, then each element of the data points is a feature. If the collected data points are image data, then each pixel or each group of pixels is a feature.

[0065] Once the input data points are manipulated or passed through gradient checks, the quantification of importance (i.e., the determination of feature importance values) is typically related to the loss or variance of the predictions. When searching for a unified measure of how good the explanation is, one cannot rely on its ability to reproduce the predictions of a machine learning model, because some methods, such as SHAP (see https: / / www.researchgate.net / publication / 317062430_A_Unified_Approach_to_interpreting_Model_Predictions), define an exact additive decomposition intended to give the predicted values. Therefore, there exists a situation where we have two additive interpretive models. and Both perfectly replicated the data instance. Prediction: .

[0066] Therefore, the two explanatory models differ in their predictions based on the contributions of the input features. The additive decomposition approach represents a "perfect" interpretation (i.e., feature importance values). However, it is entirely possible... And the corresponding interpretations may be completely different. For example, if and Then explain the model This implies that feature 1 makes a positive contribution to the prediction, while This indicates that feature 1 contributes nothing. This is a highly probable scenario, as many of these explanatory models come with mechanisms that force the explanation of sparsity.

[0067] In this example, a simple way to form meta-explanatory values ​​would be to form simple averages (such as those used by Rieger and Hansen (2020: Aggregating explanation methods for stable and robust explainability); https: / / arxiv.org / abs / 1903.00519 As described in [the document], the meta-interpretation is represented by a simple average: .

[0068] In this example, this will result in This is an unfortunate result because it will The importance of the assignment is halved and the importance of features that were not discovered at all is attributed to... Therefore, this is extremely unsatisfactory.

[0069] To provide a suitable meta-explanation, it is necessary to find aggregation rules that are more rigorously based on the original data and the behavior of the explanatory model. Therefore, a more optimized weighting rule is disclosed, which works even without accessing the ground-truth labels.

[0070] Meta-interpretation method A meta-interpreter, or meta-interpretive value, is proposed. It is a weighted combination of interpretations, specifically a weighted combination of feature importance values ​​output by different interpretation models. Therefore, the meta-interpretive value... For two or more explanatory models Each number represents a given combination, namely: , The p-th explanation model It is by Represented, and for each of the k features Output feature importance values Based on features model , for data points Provide meta-interpretive values ​​derived from the interpretive model p In a simplified way, these explanations are defined as The explanation model is additive, which makes... .

[0071] In other words, for the first The meta-explanatory value of a feature is defined as a weighted combination of the importance values ​​of the r features for that feature. Weighted pass To give, the Indicates the first Features The weights of each feature importance value.

[0072] The current goal is to address marginal constraints. and Find the optimal weight .

[0073] This paper proposes to infer these weights by measuring the sensitivity of these feature attributes as a measure of one option for a random permutation in the input space and / or a second option for a random permutation of the characteristics of the explanatory model (i.e., its hyperparameters). Furthermore, the weights can be inferred from combinations of the two options. To this end, the estimator is estimated using the following steps. variance: a) Take a data point and for all explanatory models Calculate for a given feature The corresponding feature interpretation values: .

[0074] b) Now random sampling is close to Or to Data points within the small L-norm distance This can be easily achieved using standard software packages. In this step, the hyperparameters of the interpretation method under consideration can be varied, if applicable.

[0075] c) For each sampled data point Calculate the corresponding explanation .

[0076] d) Perform the above step B times to provide one matrix , of which elements This refers to data instances. and Explanation Model and features The first explanation b One simulated value.

[0077] The above program defines a bootstrapping-like method that ideally represents the stability of the interpretation, i.e., the feature importance values ​​relative to small changes in the input data space and different hyperparameter settings. Weights corresponding to the distances from the original data points are used. The enumerative empirical variance estimator (e.g., Gaussian kernel) provides an estimate of the variance for the explanatory model: .

[0078] Given that the given explanations should be robust to small perturbations and / or hyperparameter variations, we hope to aggregate these explanations by weighting them with inverse variables: .

[0079] value Indicates a specific data collection point The weighted average of the importance values ​​of individual features in different explanatory models.

[0080] Note that, based on the reasoning above, this estimate is not only reasonable, but can also be demonstrated by standard statistical evidence that this estimator makes the overall weighted average... This minimizes the variance and therefore provides the most stable estimate for the meta-explanatory values. Furthermore, The minimum variance can even be estimated using the following calculation: , It can be calculated directly because It can be directly calculated as the result of the above procedure. This part requires and covariance The knowledge, but can be obtained from The obtained empirical covariance is directly calculated, which further allows for the identification of redundant explanatory models by examining highly correlated models.

[0081] Because the above procedure enables the creation of meta-interpretive values ​​at the local interpretation level, standard aggregation methods (such as across all data points) can be used to achieve this. Averaging the (absolute) values ​​and aggregating them into a global interpretation is trivial. Furthermore, the same logic allows for the combination of global XAI methods, i.e., a global interpretation model or even local models aggregated with a global interpretation model.

[0082] This method demonstrates how to aggregate multiple explanatory methods in a unified manner, with the sole constraint that the explanations must reference the same explanatory model and features. The disclosed method is flexible in that it allows, in principle, explanations from different categories of explanatory models, and allows for comparison of local and global explanatory models. Furthermore, as argued above, these meta-explanatory values ​​are more reliable in that they express the minimum variance explanation—that is, feature importance values ​​derived from the aggregation—and within this framework, they provide innovative measures of uncertainty. By incorporating variances arising from perturbations in the input and hyperparameter spaces, the solution should be able to integrate cognitive and stochastic uncertainties associated with computing a single explanatory method.

[0083] It is important to understand that the description of the examples above is for illustrative purposes, and the components illustrated are easily modified in various ways. For example, the concept of the illustrations can be applied to different technical systems, and especially to corresponding technical systems of different subtypes with only minor modifications.

Claims

1. A computer-implemented method for providing robust root causes of analysis results (Yi) output by an AI-based analysis model (f) of a monitoring machine (10) or the product of said machine (10), said method comprising: - Receive (S1) at least one collected data point (xi), the at least one collected data point (xi) comprising multivariable sensor data collected at the machine (10) or the product of the machine (10), each collected data point (xi) being constructed as a multidimensional vector, and each element of the vector representing a feature of the machine (10). - The analysis results (S2) are determined by processing the collected data points (xi) by the analysis model (f). - At least one meta-explanatory value (S3) indicating at least one root cause of the identified analytical result (Yi) is determined by weighted aggregation of feature importance values ​​output by at least two different explanatory models (e1, ..., en), wherein the weight values ​​depend on the stability of the feature importance value at the collected data point (xi) for each of the explanatory models (e1, ..., en), and - Output the analysis result (Yi) and the meta-explanation value (mq) representing the root cause through the user interface (S4), wherein a message M, preferably an alarm message or an instruction message, is derived from the analysis result (Yi) and the meta-explanation value (mq), and the message M is transmitted to the machine (10).

2. The computer-implemented method according to claim 1, wherein, - By inputting the analysis model (f) and the collected data points (xi) into the interpretation model (ek), the feature importance value of each feature of the collected data points (xi) to the analysis result (Yi) is determined by each of the at least two different interpretation models (e1, ..., en) (S31).

3. The computer-implemented method according to any one of the preceding claims, wherein, The weight values ​​of the explanatory model (e1, ..., en) are determined by combining the feature importance values ​​output by the explanatory model (e1, ..., en) when the input is close to the collected data point (xi). (S32) 4. The computer-implemented method according to any one of the preceding claims, wherein, The weight values ​​of the explanatory model (e1, ..., en) are determined (S33) by combining the feature importance values ​​output by the modified version of the explanatory model (e1, ..., en), which differs through the hyperparameters of the explanatory model (e1, ..., en).

5. The computer-implemented method according to any one of the preceding claims, wherein, Redundant explanatory models are identified by determining highly relevant explanatory models, where the redundant explanatory models are no longer used to determine feature importance values.

6. The computer-implemented method according to any one of the preceding claims, wherein, The global meta-interpretation value for all data points (xi) is generated by an aggregation method over all data points (xi), preferably by averaging the weight values ​​over all data points (xi).

7. The computer-implemented method according to any one of the preceding claims, wherein, The interpretation model (el, ...en) is either a local interpretation model or a global interpretation model.

8. The computer-implemented method according to any one of the preceding claims, wherein, The local interpretation model and / or the global interpretation model are model-agnostic.

9. The computer-implemented method according to any one of the preceding claims, wherein, The analytical model (f) is a machine learning model configured to perform classification or regression tasks.

10. The method according to any one of the preceding claims, wherein, The analysis task is one of anomaly detection, predictive maintenance prediction, and quality monitoring.

11. An apparatus (20) for analyzing a machine (10) or a product of said machine, said apparatus (20) comprising: - Data interface (21), configured to receive at least one collected data point (xi), the at least one collected data point (xi) comprising multivariable sensor data collected at the machine (10) or the product of the machine (10), each collected data point (xi) being constructed as a multidimensional vector, and each element of the vector representing a feature of the machine (10), -Analyzer unit (22), which is configured to determine the analysis result (Yi) about the analysis task by processing the collected data points (xi) by the analysis model (f). - An interpreter unit (23) is configured to determine a meta-interpretive value indicating at least one root cause of the determined analytical result (Yi) by weighted aggregation of feature importance values ​​output by at least two different interpretive models (e1, ..., en), wherein the weight values ​​depend on the stability of the feature importance value at the collected data point (xi) for each of the interpretive models (e1, ..., en), and - User interface (24), which is configured to output the analysis results (Yi) and the meta-explanation value (mq) representing the root cause. Includes a machine interface (25) configured to derive a message (M), preferably an alarm message or an instruction message, from the analysis result (Yi) and the meta-interpretation value (mq) and transmit the message (M) to the machine (10).

12. The apparatus according to any one of the preceding claims, wherein, The device (20) is one of an anomaly detector, a predictive maintenance evaluator, and a quality monitor.

13. A computer program product capable of being directly loaded into the internal memory of a digital computer, the computer program product comprising software code portions for performing the steps of claims 1 to 10 when the product is run on the digital computer.