Ranking of causal anomalies via temporal and dynamic analysis of vanishing correlations

DE112017000704B4Active Publication Date: 2026-08-27NEC CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
DE112017000704
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2017-02-01
Filing Date
2017-02-01
Publication Date
2026-08-27
Estimated Expiration
2037-02-01

AI Technical Summary

Technical Problem

Existing methods for detecting anomalies in large-scale information systems fail to effectively identify causal anomalies and diagnose system failures due to the lack of consideration for anomaly propagation and temporal dynamics in invariant networks.

Method used

A computer-implemented method using a random walk-based framework to model anomaly propagation in invariant networks, restore broken links, and perform temporal smoothing to rank causal anomalies, incorporating a sparse penalty and optimization to identify root causes.

Benefits of technology

Enhances the accuracy of detecting and diagnosing system failures by identifying the most likely causal anomalies, reducing the effort required for debugging and maintenance in complex systems.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

A computer-implemented method for root cause anomaly detection in a multi-node invariant network generating time-series data, comprising: modeling anomaly propagation in the invariant network by a processor; restoring broken invariant links in an invariant graph based on rank-order vectors of causal anomalies by the processor, wherein each of the broken invariant links comprises a respective pair of nodes formed from the multiple nodes such that one of the nodes in the respective pair of nodes has an anomaly, wherein each of the rank-order vectors of causal anomalies serves to specify, for a given set of the multiple nodes when arranged in pairs, a respective node anomaly status; computing a sparse penalty of the rank-order vectors of causal anomalies to obtain a set of time-dependent anomaly orders by the processor;Performing time-smoothing of the set of time-dependent anomaly orders by the processor; and controlling an anomaly initiation by one of the multiple nodes based on the set of time-dependent anomaly orders by the processor.
Need to check novelty before this filing date? Find Prior Art

Description

INFORMATION ABOUT RELATIVES REGISTRATION

[0001] This application claims priority over the preliminary US patent application serial no. 62 / 292383, filed on February 8, 2016, which is incorporated herein in its entirety by reference. BACKGROUND Technical area

[0002] The present invention relates to computer learning and in particular to the ranking of causal anomalies via temporal and dynamic analysis of vanishing correlations. Description of the related area

[0003] With the rapid advances in networking and computer technology, the complexity of networked applications and information services is increasing dramatically. These large-scale information systems typically contain thousands of components. Therefore, there is a need for these systems to automatically monitor system status, detect anomalies, and diagnose system malfunctions. This is crucial for enabling human decision-making in equipment and system maintenance and troubleshooting. SUMMARY

[0004] According to one aspect of the present invention, a computer-implemented method for root cause anomaly detection in a multi-node invariant network generating time-series data is provided. The method includes modeling anomaly propagation in the invariant network by a processor. Furthermore, the method includes restoring broken invariant connections in an invariant graph based on rank-order vectors of causal anomalies by the processor. Each broken invariant connection comprises a respective pair of nodes formed from the multiple nodes such that one of the nodes in the respective pair exhibits an anomaly. Each of the rank-order vectors of causal anomalies serves to specify a respective node anomaly status for a given set of multiple nodes when arranged in pairs.Furthermore, the procedure includes the processor calculating a sparse penalty on the rank order vectors of causal anomalies to obtain a set of time-dependent anomaly orders. Additionally, the procedure includes the processor performing time smoothing on this set of time-dependent anomaly orders. Finally, the procedure includes the processor initiating an anomaly at one of the multiple nodes based on this set of time-dependent anomaly orders.

[0005] According to another aspect of the present invention, a computer program product for root cause anomaly detection in a multi-node invariant network generating time-series data is provided. The computer program product includes a non-transitory, computer-readable storage medium on which program instructions embodied therein are stored. The program instructions are executable by a computer to cause the computer to execute a method. The method includes the modeling of anomaly propagation in the invariant network by a processor. Furthermore, the method includes the recovery of broken invariant connections in an invariant graph based on rank-order vectors of causal anomalies by the processor.Each of the broken invariant connections comprises a respective pair of nodes, formed from the multiple nodes such that one node in the respective pair exhibits an anomaly. Each of the rank-order vectors of causal anomalies serves to specify a respective node anomaly status for a given set of the multiple nodes when arranged in pairs. Furthermore, the procedure includes the processor calculating a sparse penalty of the rank-order vectors of causal anomalies to obtain a set of time-dependent anomaly orders. Additionally, the procedure includes the processor performing time smoothing of the set of time-dependent anomaly orders. Finally, the procedure includes the processor controlling the initiation of an anomaly at one of the multiple nodes based on the set of time-dependent anomaly orders.

[0006] According to yet another aspect of the present invention, a computer processing system for root cause anomaly detection in a multi-node invariant network generating time-series data is provided. The system includes a processor. The processor is configured to model anomaly propagation in the invariant network. Furthermore, the processor is configured to reconstruct broken invariant connections in an invariant graph based on rank-order vectors of causal anomalies. Each broken invariant connection comprises a pair of nodes formed from the multiple nodes such that one node in the respective pair exhibits an anomaly. Each rank-order vector of causal anomalies serves to specify a respective node anomaly status for a given set of multiple nodes when arranged in pairs.Furthermore, the processor is configured to compute a sparse penalty on the rank order vectors of causal anomalies to obtain a set of time-dependent anomaly orders. Additionally, the processor is configured to perform time smoothing on this set of time-dependent anomaly orders. Finally, the processor is configured to control anomaly initiation at one of the multiple nodes based on this set of time-dependent anomaly orders.

[0007] These and other features and advantages will become apparent from the following detailed description of illustrative embodiments, which should be read in conjunction with the accompanying drawings. List of characters

[0008] The disclosure provides details of preferred embodiments in the following description with reference to the following figures, in which: Fig.Figure 1 shows a block diagram of an exemplary processing system 100 to which the present invention can be applied, according to an embodiment of the present invention; Fig. 2 shows a block diagram of an exemplary environment 200 to which the present invention can be applied, according to an embodiment of the present invention; Fig. Figure 3 shows a high-level block diagram / flowchart of an exemplary ranking system / ranking procedure for 300 causal anomalies according to an embodiment of the present invention; and Fig. Figure 4 shows a flow chart of an exemplary procedure 400 for ranking causal anomalies according to an embodiment of the present invention. DETAILED DESCRIPTION OF PREFERRED EXECUTION FORMS

[0009] The present invention is directed towards the ranking of causal anomalies via temporal and dynamic analysis of vanishing correlations.

[0010] Detecting anomalies in monitoring data from distributed information systems and cybernetic-physical systems is an essential task with numerous applications in fields such as industry, security, and healthcare. Discovering invariant relationships between monitoring data and generating invariant networks has been shown to be an effective way to characterize relationships between system components. In an invariant network, a node represents a unit of monitoring data, and a link indicates a correlation between two units of monitoring data. Such an invariant network can help systems experts discover anomalies and diagnose system malfunctions by examining these vanishing correlations.

[0011] In one embodiment of the present invention, a random-walk-based (or network propagation-based) framework for the ranking of causal anomalies is proposed to identify important causal anomalies in the invariant network and to provide their relative importance with respect to the probability of being the root cause of the system disturbance.

[0012] First, the invariance graph and the fractional pairs of the invariance graph are used as input. Then, the random walk model is used to estimate the propagation of the true causal anomalies in the invariance graph. Afterward, an optimization framework is used that has the following two terms: ( 1 ) Minimizing the number of true anomalies in the graph, since the system anomaly or system disturbance, based on prior knowledge, is usually triggered by a limited number of "seed" anomalies; and ( 2The propagated anomalies from the seeds should be consistent with the broken edges provided in the invariance graph by minimizing their recovery error. The inventors designed an iterative optimization procedure to obtain a locally optimal solution.

[0013] Fig. Figure 1 shows a block diagram of an example processing system. 100 , to which the inventive principles can be applied, according to an embodiment of the present invention. The processing system 100 contains at least one processor (a CPU) 104 , which uses a bus system 102 It is functionally coupled with other components. With the bus system 102 are a cache 106 , a read-only memory (ROM) 108 , a read / write memory (RAM) 110 , an input / output adapter (I / O adapter) 120 , a sound adapter 130 , a power adapter 140, a user interface adapter 150 and a display adapter 160 functionally coupled.

[0014] With the bus system 102 are through the I / O adapter 120 a first storage device 122 and a second storage device 124 functionally coupled. The storage devices 122 and 124 These can be a disk storage device (e.g., a magnetic or optical disk storage device) and / or a magnetic solid-state device, etc. The storage devices 122 and 124 They can be of the same storage device type or of different storage device types.

[0015] With the system bus 102 is through the sound adapter 130 a loudspeaker 132 functionally coupled. With the system bus 102 is through the power adapter 140 a transceiver 142functionally coupled. With the system bus 102 is through the display adapter 160 a display device 162 functionally coupled.

[0016] With the system bus 102 are through the user interface adapter 150 a first user input device 152 , a second user input device 154 and a third user input device 156 functionally coupled. The user input devices 152 , 154 and 156 These can be a keyboard, a mouse, a keypad, an image capture device, a motion capture device, a microphone, or a device that incorporates the functionality of at least two of the preceding devices, etc. Of course, other types of input devices can also be used while maintaining the inventive concept of the present principles. The user input devices152 , 154 and 156 They can be of the same user input device type or of different user input device types. The user input devices 152 , 154 and 156 are used for entering and outputting information into and out of the system 100 used.

[0017] As any expert in the field can easily understand, the processing system 100 Naturally, it may also contain other elements (not shown), and certain elements may also be omitted. As the average expert in the field can easily understand, depending on the specific implementation, various other input and / or output devices may be included in the processing system. 100It may include, for example, various types of wireless and / or wired input and / or output devices. As the person skilled in the art will readily appreciate, additional processors, controllers, memory, etc., in various configurations may also be used. The person skilled in the art, using the teachings of the present invention, will be able to devise these and other variants of the processing system. 100 Easier said than done.

[0018] Furthermore, it will be acknowledged that the following will be based on Fig. 2 described environment 200 An environment for implementing respective embodiments of the present invention. Part of the processing system 100 or the entire processing system can be located in one or more of the elements of the environment 200 be implemented.

[0019] Furthermore, it will be acknowledged that the following will be based on Fig. 3 described systems 300 A system for implementing respective embodiments of the present invention. Part of the processing system 100 or the entire processing system can be in one or more of the elements of the system 300 be implemented.

[0020] Furthermore, it will be acknowledged that the processing system 100 at least a part of the procedure described here, including, for example, at least a part of the procedure 400 out of Fig. 4 can execute. Similarly, part of the environment can 200 or the entire environment is used to carry out at least part of the process 400 out of Fig. 4. To execute. Additionally, part of the system can 300 or the entire system can be used to perform at least part of the process 400 out of Fig. 4 to execute.

[0021] Fig. Figure 2 shows a block diagram of an example environment 200 , to which the present invention can be applied, according to one embodiment of the present invention. The environment 200 represents an invariant computer network to which the present invention can be applied. The information relating to Fig. The two elements shown are for illustrative purposes only. However, it will be appreciated that, as the person skilled in the art in the field will readily understand from the teachings given herein, the present invention can be applied to other network configurations while preserving the inventive concept of the present invention.

[0022] The environment 200 contains at least a set of nodes that can be identified individually and collectively by the figure reference symbol 210 are designated. Each of the nodes 210may include one or more servers or other types of computer processing equipment, individually and collectively identified by the figure reference symbol 211 are designated. The computer processing devices 211 They can contain, for example, machines (e.g., industrial machines, assembly line machines, robots, etc.), but are not limited to them. For illustration, each of the nodes 210 with a large number of servers 211 shown. Each of the nodes generates time series data and / or makes it available in another way.

[0023] As described herein, in one embodiment of the present invention, causal anomalies in the network are ranked by means of temporal and dynamic analysis of vanishing correlations. Based on these rankings, a computer processing system can be controlled to mitigate errors resulting from the propagation of a causal anomaly.

[0024] In the Fig. In the embodiment shown in 2, the elements thereof are divided by one or more nets. 201 interconnected. However, other types of connections can also be used in other embodiments. Additionally, one or more elements can be incorporated into Fig. 2. can be implemented by a variety of devices that include, but are not limited to, digital signal processing circuits (DSP circuits), programmable processors, application-specific integrated circuits (ASICs), free-programmable logic arrays (FPGAs), complex programmable logic devices (CPLDs), etc. These and other variants of the environment elements 200 are easily determined by the person skilled in the art in the field using the teachings of the present invention given herein, while preserving the inventive concept of the present invention.

[0025] Fig.Figure 3 shows a high-level block diagram / flowchart of an exemplary ranking system / ranking procedure. 300 causal anomalies according to an embodiment of the present invention.

[0026] The system / procedure 300 contains a set of single-time ranking devices 310 and an institution 320 for time smoothing. Each of the single-time ranking devices 310 contains an invariance graph 311 , couples 312 broken invariance, a random-walk reproductive device 313 , a recovery error detection device 314 and an optimization device 315 .

[0027] Each of the single-time ranking facilities 310 performs a ranking at a single point in time. That is, each output of the set of single-point ranking devices. 311corresponds to a specific point in time (e.g., the result at time t, the result at time t + 1, ..., the result at time T).

[0028] The Radom-Walk Reproductive Facility 313 performs a random walk with restart to trace the propagation process from a few germline abnormalities to the entire invariance graph. 311 to model.

[0029] The recovery error detection device 314 models the recovery error of the propagated anomalies and the pairs 312 broken invariance.

[0030] The optimization facility 315 executes an iterative optimization algorithm to compute the sparse causal anomaly vector.

[0031] The facility 320Time smoothing is used to force the smoothness of the ranking result at adjacent time points in order to improve the global consistency of the ranking results.

[0032] Furthermore, regarding the Random Walk Reproduction Facility 313 A random walk restart (RWR) technique is used to model potential propagation. Assuming that e denotes the indicator vector in which ei indicates whether the corresponding node in the invariant network is a causal anomaly, the corresponding entry ei is set to 1 for all nodes of a causal anomaly, and all other entries are set to 1. 0 set. Subsequently, the anomalous state of e propagates other nodes with the following objective function: min r ≥ 0 c ∑ i , j=1 n A ij ‖ 1 D ii r i − 1 D jj r j ‖ 2 + ( 1 − c ) ∑ i = 1 n ‖ r i − e i ‖ 2 min r ≥ 0 c r T ( I n − A ˜ ) r+ ( 1 − c ) ‖ r − e ‖ F . 2

[0033] Here, matrix A is the adjacency of the invariant network, D is the diagonal degree matrix of A, c is a scalar controlling the weight between the strength of propagation, r is the propagation outcome vector, and e are the initial seeds of anomalies that are actually responsible for subsequent fractional invariances. The convergent solution of r can be written as follows: r = ( 1 − c ) ( I n − c A ˜ ) − 1 e .

[0034] Furthermore, regarding the recovery error determination device 314 the initial germ vector e, after it has propagated in the invariance network, to the pairs 312broken invariance; that is, because under anomalous conditions, two previously correlated nodes (or an invariance) are broken if an anomaly occurs, meaning that one or both of the nodes have intrinsic changes in their state and the two nodes are therefore no longer synchronized with each other. The pairs 312 Broken invariances have been recorded in the graph whose adjacency matrix is ​​P̃. If the i-th and j-th nodes are no longer correlated, the corresponding entry in P̃ is set to 1; otherwise, it is zero. The recovery error is measured as follows: min e i ∈ { 0,1 } ,1 ≤ i ≤ n ‖ ( Bee T B T ) ∘ M − P ˜ ‖ P F 2 .

[0035] Furthermore, regarding the optimization device 315 by combining the Random Walk Reproduction Facility 313 and the recovery error detection device 314 A global optimization framework would be as follows: min e ≥ 0 ‖ ( Bee T B T ) ∘ M − P ˜ ‖ F 2 + τ ‖ e ‖ 1 .

[0036] The optimization variable here is e, and there are two terms. The first term requires that the propagated anomaly should be consistent with the fractional invariance (i.e., recovery of the fractional pairs); the second term is a penalty term on the 11-norm of the variable e, which promotes many zero entries in the vector e. This optimization problem can be computed using the following iteration: e ← e ∘ { 4 [ ( B T P ˜ ) ∘ M ] Be 4 [ ( B T Bee T B T ) ∘ M ] Be + τ 1 n } 1 4 .

[0037] Furthermore, regarding the facility 320For temporal smoothing in anomaly detection, it is a reasonable assumption that the anomalies propagate in the invariant network as time passes. However, the underlying causal anomalies typically remain unchanged within a time interval T. Based on this intuition, a smoothing procedure is developed here that considers temporally and dynamically broken networks together. That is, a smoothing term is added to the objective functions. Here, e(i-1) and e(i) are the rank-order vectors of causal anomalies at two continuous time points. The objective function can then be written as follows: min e ( i ) ≥ 0,1 ≤ i ≤ T ∑ i = 1 T [ ‖ ( Be ( i ) ( e ( i ) ) T B T ) ∘ M − P ˜ ‖ F 2 + τ ‖ e ( i ) ‖ 1 ] + α ‖ e ( i ) − e ( i − 1 ) ‖ 2 2 .

[0038] An iterative optimization procedure can be defined as follows: e ( i ) ← e ( i ) ∘ { 4 [ ( B T P ˜ ) ∘ M ] Be+2 α e ( i − 1 ) 4 [ ( B T Bee T B T ) ∘ M ] Be + τ 1 n +2 α e ( i ) } 1 4 .

[0039] A description of conventional methods and some of the differences between the conventional methods and the present invention will now be given.

[0040] Conventional methods typically use neural networks to estimate the input and output mapping functions, which requires tuning many parameters and can lead to a locally optimal solution. Here, a manifold-regularized kernel regression framework is used, which can provide a globally optimal solution and makes it more practical to locate the control parameter in the optimal KPI values. Another difference is the use of a data-driven mathematical procedure to separate variables that are not well explained by the input variables, which can improve online optimization. Conventional methods typically do not identify such variables from the input data.

[0041] Fig. Figure 4 shows a flowchart of an example procedure. 400for the ranking of causal anomalies according to an embodiment of the present invention.

[0042] In step 405 Anomaly propagation is modeled in the invariant network. In one embodiment, anomaly propagation in the invariant network is modeled using a random walk with restart technique. In another embodiment, anomaly propagation in the invariant network is modeled after the propagation of an initial perturbation in the rank-order vectors of causal anomalies using an objective function based on an anomaly weight vector. In yet another embodiment, anomaly propagation in the invariant network is modeled based on a threshold applied to a triple consisting of a first-degree autoregressive-eXogenous (AXR) model, a second-degree AXR model, and a time delay between the time-series data.

[0043] In step 410 Broken invariant connections are reconstructed in an invariant graph based on rank-order vectors of causal anomalies. Each broken invariant connection contains a corresponding pair of nodes formed from the multiple nodes such that one node in the respective pair exhibits an anomaly. Each rank-order vector of causal anomalies, when arranged in pairs, specifies a given node anomaly status for a given set of nodes. In one embodiment, the broken invariant connections are identified based on an objective function.

[0044] In step 415 A recovery error for the invariant network is determined based on the broken invariant links.

[0045] In step 420The set of time-dependent anomaly orders is determined based on the recovery error.

[0046] In step 425 A sparse penalty of the rank order vectors of causal anomalies is calculated to obtain a set of time-dependent anomaly orders. In one embodiment, the sparse penalty of the rank order vectors of causal anomalies is used to control the number of non-zero values ​​in the set of time-dependent anomaly orders.

[0047] In step 430The set of time-dependent anomaly orders is optimized. In one embodiment, the set of time-dependent anomaly orders is optimized using an objective function. In one embodiment, the objective function has (i) a first term that requires consistency between a propagated anomaly and a broken invariant combination, and (ii) a second term that is a penalty term to encourage zero entries in the rank order vectors of causal anomalies.

[0048] In step 435 A time-smoothing process is performed on the set of time-dependent anomaly arrays.

[0049] In step 440The initiation of an anomaly at one of the multiple nodes is controlled based on the set of time-dependent anomaly arrays. In one embodiment, the control can include switching off a root cause computer processing device upon anomaly initiation at one of the multiple nodes in order to mitigate error propagation. In another embodiment, the control can include terminating a root cause process that is running in a computer processing device at the node initiating an anomaly, in order to mitigate error propagation.

[0050] A description regarding system invariants and vanishing correlations according to an embodiment of the present invention is now given.

[0051] A basic framework for detecting pairwise correlations in massive time series is described. Correlations in a normal system phase are referred to as system invariants. Each correlation is called an invariant link. Subsequently, the method for detecting vanishing correlations during an anomalous phase of the system is described. These vanishing correlations are also referred to here as broken invariant links.

[0052] An invariant model according to an embodiment of the present invention will now be described.

[0053] An invariant is a model that describes a pairwise relationship between time series, expressed as an autoregressive-eXogenous (ARX) model that accounts for the time lag. Let x(t) and y(t) be the observed values ​​from the time series x and y, respectively, at time t, let n and m be the degrees of the ARX model, and let k be the time lag. Let ŷ(t;θ) be the estimate of y(t) using the ARX model parameterized by θ. It can be expressed as follows: y ^ ( t ; θ ) = a 1 y ( t − 1 ) + ⋯ + a n y ( t − n ) + b o x ( t − k ) + ⋯ + b m x ( t − k − m ) + d ( 1 ) = φ ( t ) T θ , ( 2 ) where θ= [a1, ..., a n , b0, ..., b m , d] T ∈ ℝ n+m+2 , φ(t) = [y(t-1), ...,y(tn),x(tk),...,x(tkm), 1] T ∈ ℝ n+m+2 are. Table 1: Summary of spellings symbol definition n the number of nodes in the invariant graph c, λ, τ the parameters 0 < c < 1, τ > 0, λ > 0 σ(▪) the Softmax function G l the invariant network G b the broken network for G l A (Ã) ∈ ℝ nxn the (normalized) adjacency matrix of G l P (P̃) ∈ ℝ nxn the (normalized) adjacency matrix of G b M ∈ ∈ ℝ nxn the logical matrix of G l d(i) the degree of the i-th node in graph G l D ∈ ℝ nxn the degree matrix: D = diag(d(i), ..., d(n)) r ∈ ℝ nx1 the anomaly weight vector e ∈ ℝ nx1 the rank order vector of causal anomalies

[0054] For a fixed (n, m, k), the parameter θ can be estimated by least squares at all observed time points t = 1, ..., N in the training time series. In practice, only 'good' ARX models should be used for anomaly detection, so the 'goodness of fit' of an ARX model is defined. The following uses a suitability assessment F(θ), which is defined as follows: F ( θ ) = 1 − ∑ t = 1 N | y ( t ) y ^ ( t ; θ ) | 2 ∑ t = 1 N | y ( t ) - y ¯ ) | 2 , where y̅ is the average of all observed values ​​y(t). F(θ) is always less than 1, and a higher F(θ) indicates that the ARX model fits the observed data well. If a threshold is specified, and if the fit rating of an ARX model for x and y is greater than the threshold, it is declared that there is an invariant (correlation) between them. The network containing all invariant links that encode the pairwise correlations is called the invariant network. This procedure of constructing a system-invariant network is referred to here as the model training time. The inferred θ is used to dynamically track vanishing correlations between each pair of time series during the training time.

[0055] A description of the detection of vanishing correlations according to an embodiment of the present invention is now given.

[0056] To detect vanishing correlations during an anomalous system phase, the simplest approach for ARX model selection is to consider all possible combinations of (m, n, k) within a prefixed range and select the model with the highest suitability score. In real-world applications such as anomaly detection in physical systems, 0 ≤ n, m, k ≤ 2 is commonly used.

[0057] The invariant constructed using the method described here is used to track vanishing correlations in real time via the following procedure. At each time point, the (normalized) remainder R(t) between the measured value y(t) and its estimated value ŷ(t; θ), defined as follows, is calculated: R ( t ) = | y ( t ) − y ^ ( t ; θ ) | ε max , where ε max the maximum error in training ARX models, i.e. ε max max1 ≤ t < n|y (t)- ŷ(t; θ)|, is. If the remainder exceeds a predetermined threshold, the invariant is declared 'broken', meaning the correlation between the two time series vanishes. The network containing all vanishing correlations of the system is called the broken network. This procedure of tracking the broken network of the system is called the testing period.

[0058] A description of a problem addressed by the present invention will now be given.

[0059] It is G l the invariant network (the graph 2 ), which contains n nodes, and let G b the broken network for G l To determine the adjacency matrix of network G l or G b To denote, two symmetric matrices A ∈ ℝ are used. nxn , P ∈ ℝ nxnThese two networks can be obtained using the techniques introduced here. The values ​​of the two matrices can be either binary or real values. In the binary case, 1 is used to indicate that there is a correlation between two corresponding time series, while 0 means no correlation. In the case of real values, for example, the suitability rating F(θ) or the remainder R(t) can be used as the values ​​of the two matrices.

[0060] The goal is to causally identify anomalous nodes in G l to detect those most likely to have the broken status in G b The present invention provides effective algorithms for assigning a ranking to the nodes such that the nodes with the highest ranking are most likely to be the causal anomalies. The ranking vector of the causal anomalies is denoted as e. Important notations are listed in Table 1.

[0061] A description of the algorithm for ranking causal anomalies according to an embodiment of the present invention is now given.

[0062] In particular, an algorithm for the ranking of causal anomalies (RCA) is described. The ranking of causal anomalies is modeled as the recovery error minimization problem. The proposed RCA method simultaneously optimizes the empirical probability of the broken network and also takes into account error propagation in the invariant network.

[0063] A description of an objective function used by one or more embodiments of the present invention will now be given.

[0064] Due to various system error behaviors, system noise, data uncertainties, etc., it is impractical to accurately determine the causal anomalies at the earliest time point. To mitigate this problem, the random walk restart (RWR) technique is used to model error propagation in the invariant network. It is assumed that e denotes the indicator vector in which e i (1 ≤ i ≤ n) indicates whether the corresponding node in the invariant network is a causal anomaly. The entry e i will be applied to all nodes of causal anomalies 1 set and otherwise 0. The anomalous state e then propagates to other nodes with the following objective function: min r ≥ 0 c ∑ i , j=1 n A ij ‖ 1 D ii r i − 1 D jj r j ‖ 2 + ( 1 − c ) ∑ i = 1 n ‖ r i − e i ‖ 2 , where D ∈ ℝ nxnThe degree matrix of A is c ∈ (0, 1), the regularization parameter is c, and r is the anomaly weight vector (anomaly evaluation vector) after propagation of the initial perturbation into e. Equation 5 is equivalent to the following formula: min r ≥ 0 c r T ( I n − A ˜ ) r+ ( 1 − c ) ‖ r − e ‖ F 2 , where à is the normalized A and equal to D -½ AD -½Similarly, the normalization of P is denoted as P̃. The first term in Equation 6 is the smoothness boundary condition, which means that a good rank order function should not change too much between nearby points in the invariant network. The second term is the fitting boundary condition, which means that a good rank order function should not deviate too much from the initial anomaly assignment. The trade-off between these two competing boundary conditions is controlled by a positive parameter c. Since α is stochastic, Equation 6 converges for a stationary point r using the following formula: r = ( 1 − c ) ( I n − c A ˜ ) − 1 e .

[0065] To encode the information of a broken network, r is used to reconstruct the broken network. If it is in G b a broken operation exists, e.g. P̃ ijIf ≠ 0, ideally at least one of nodes i and j is anomalous. It is noted that this is not necessarily the causal anomaly, but rather that it could be the error in the information flow direction behind it, caused by potential causal anomalies. For this purpose, either r i or r j be large. Thus, the product of r i and r j used to determine the value of P̃ ij to restore it. The following describes a procedure to normalize it in order to avoid extreme values. Subsequently, the loss of the restoration of the fractional link P̃ can be corrected. ij by (r i ▪r j - P̃ ij ) 2 be calculated. Thus, the recovery error is for the entire broken network. ‖ ( rr T ) ∘ M − P ˜ ‖ F 2 . Here, ∘ is an element-wise operator and M is the logical matrix of the invariant network G. l(1 with edge, 0 without edge). Let B = (1 - c)(I n - cÃ) -1 , where substituting r into the recovery error formula yields the following objective function: min e i ∈ { 0,1 } ,1 ≤ i ≤ n ‖ ( Bee T B T ) ∘ M − P ˜ ‖ F 2 .

[0066] Considering that the integer programming in Equation 8 is difficult to solve and that the goal here is to assign the ranking of the causal anomalies, this is relaxed using the penalty ℓ1 on e with parameter τ to control the number of non-zero entries in e. This results in the following objective function: min e ≥ 0 ‖ ( Bee T B T ) ∘ M − P ˜ ‖ F 2 + τ ‖ e ‖ 1 .

[0067] A description of a learning algorithm according to an embodiment of the present invention is now given.

[0068] An iterative multiplicative update algorithm is provided here to optimize the objective function in Equation 9. The objective function is invariant under these updates if and only if e is a stationary point. More precisely, the solution to the optimization problem in Equation 9 relies on the following theorem, which is derived from the Karush-Kuhn-Tucker complementarity condition (KKT complementarity condition).

[0069] Sentence 1 Updating e according to equation 10 reduces the objective function in equation 9 monotonically until convergence. e ← e ∘ { 4 [ ( B T P ˜ ) ∘ M ] Be 4 [ ( B T Bee T B T ) ∘ M ] Be + τ 1 n } 1 4 , where ∘ , [ ⋅ ] [ ⋅ ] und ( ⋅ ) 1 4 These are element-wise operators.

[0070] Based on sentence 1 Here, the iterative multiplicative update algorithm is used for its optimization and summation in the algorithm. 1 developed. The ranking algorithm is referred to here as RCA.

[0071] A description of a theoretical analysis according to an embodiment of the present invention will now be given.

[0072] Initially, the derivation is described as follows.

[0073] The solution to equation (10) is derived according to the theory of optimization under given boundary conditions. Since the objective function is not jointly convex, an efficient multiplicative update algorithm is assumed for the optimization to find a locally optimal solution. The theorem 1 This will be proven as follows. The Lagrange function will be used for the optimization of L = ‖ ( Bee T B T ) ∘ M − P ˜ ‖ F 2 + τ 1 n T e <?page 13=""?> formulated. Obviously, B, M, and P̃ are symmetric matrices. Let F = (Bee T B T )◦M, where the following then applies: ∂ ∂ e m ( F − P ˜ ) ij 2 = 2 ( F ij − P ˜ ij ) ∂ F ij e m = 4 ( F ij − P ˜ ij ) M ij ( B mi T B j : e ) ( gemäß Symmetrie ) = 4 B mi T ( F ij − P ˜ ij ) M ij ( Be ) j : .

[0074] It follows that ∂ ‖ F − P ˜ ‖ F 2 ∂ e m = 4 B m : T [ ( F − P ˜ ) ∘ M ] ( Be ) and thereby ∂ ‖ F − P ˜ ‖ F 2 ∂ e = 4 B T [ ( F − P ˜ ) ∘ M ] ( Be ) applies.

[0075] Thus, the partial derivative of the Langrange function with respect to e is as follows: ∇ eL = 4 B T [ ( Bee T B T − P ˜ ) ∘ M ] Be + τ 1 n , where 1 n The n × 1 vector consists only of ones. Using the Karush-Kuhn-Tucker complementarity condition (KKT complementarity condition) for the non-negative boundary condition on e, the following results: ∇ eL ∘ e = 0.

[0076] The above formula leads to the update rule for e shown in equation 10.

[0077] A description of the convergence according to one embodiment of the present invention will now be given.

[0078] To verify the convergence of equation (10) in the theorem 1To prove this, the auxiliary function approach is used. The definition of the auxiliary function is introduced as follows:

[0079] Definition 4.1 Z(h,ĥ) is an auxiliary function for L(h) if for any given h, ĥ the following boundary conditions are satisfied: z ( h , h ^ ) ≥ L ( h ) und Z ( h , h ) = L ( h ) .

[0080] lemma 4.1 If Z is an auxiliary function for L, then L does not increase under the update. h ( t+1 ) = argmin h Z ( h , h ( t ) ) .

[0081] Sentence 2 Let L(e) denote the sum of all terms in L that contain e. The following function is an auxiliary function for L(e): Z ( e , e ^ ) = − 2 ∑ ij e ^ i { [ ( B T P ˜ ) ∘ M ] B } ij e ^ j ( 1 + log e i e j e ^ i e ^ j ) + ∑ i { [ ( B T B e ^ e ^ T B T ) ∘ M ] B e ^ } i e i 4 e ^ i 3 + τ 4 ∑ i e i 4 + 3 e ^ i 4 e ^ i 3 .

[0082] Furthermore, it is a convex function in e and has a global minimum.

[0083] The sentence 2 can be done by validating Z(e,ê) ≥ L(e), Z(e, e) = L(e) and the Hessian matrix ∇∇ e Z(e, ê) ≥ 0 can be proven.

[0084] Based on the sentence 2 Z(e, ê) can be minimized with respect to e, where ê is fixed. For this purpose, ∇ is used. e Z(e, ê) = 0 was set and the following update formula was obtained: e ← e ^ ∘ { 4 [ ( B T P ˜ ) ∘ M ] B e ^ 4 [ ( B T B e ^ e ^ T B T ) ∘ M ] B e ^ + τ 1 n } 1 4 , which is consistent with the update formula derived from the KKT condition mentioned above.

[0085] From Lemma 4.1 and from sentence 2 For each subsequent iteration of the update of e, the following results: L(e 0 ) = Z(e 0 , e 0 ) ≥ Z(e 1 , e 0 ) ≥ Z(e 1 , e 1 ) = L(e 1 ) ≥ ... ≥ L(e Iter ). Thus, L(e) decreases monotonically. Since the objective function equation (9) is bounded below by 0, the theorem is true. 1 proven. The sentence 1 can be proven with a similar strategy.

[0086] A complexity analysis according to one embodiment of the following invention will now be described.

[0087] In algorithm 1 The inverse of an n × n matrix must be calculated, which is associated with the complexity (n 3 ). In each iteration, multiplication between two n × n matrices is unavoidable, so the overall time complexity of the algorithm 1 (Iter ▪ n 3 ) where Iter is the number of iterations required before convergence. An alternative algorithm is proposed below that avoids computing the inverse n × n matrix and multiplication between two n × n matrices. The time complexity can be expressed as (Iter ▪n 2 ) will be reduced.

[0088] A description of a computation acceleration according to an embodiment of the present invention is now given.

[0089] The analysis described above has determined that the time complexity of the algorithm 1 (Iter ▪ n 3 ). Another algorithm is proposed that avoids calculating the inverse of the n × n matrix and multiplication between two n × n matrices. The time complexity can be reduced to (Iter ▪ n 2 The computation speed is reduced by relaxing the objective function in Equation 9 to optimize the anomaly weight vector r and the rank order vector of causal anomalies e together. The objective function is as follows: min e ≥ 0, r ≥ 0 cr T ( I n − A ˜ ) r + ( 1 − c ) ‖ r − e ‖ F 2 + λ ‖ ( rr T ) ∘ M- P ˜ ‖ F 2 + τ ‖ e ‖ 1 .

[0090] To solve the above objective function, an alternating scheme can be used. That is, the objective function is optimized with respect to r while e is fixed, and vice versa. This procedure is continued until convergence. The objective function is invariant under these updates if and only if r, e is a stationary point. More precisely, the solution to the optimization problem in Equation 20 relies on the following theorem, which is derived from the Karush-Kuhn-Tucker complementarity condition (KKT complementarity condition). Its derivation and the proof of the theorem 3 are similar to sentence 1 .

[0091] Sentence 3 . Alternating updates of e and r according to equation 21 and equation 22 decrease the objective function in equation 20 monotonically until convergence. r ← r ∘ { A ˜ r + 2 λ ( P ˜ ∘ M ) r + ( 1 − c ) e r + 2 λ [ ( rr T ) ∘ M ] r } 1 4 , e ← e ∘ [ 2 ( 1 − c ) r τ 1 n + 2 ( 1 − c ) e ] 1 2 .

[0092] Based on the sentence 3Can an iterative multiplicative update algorithm be used for optimization, similar to the algorithm? 1 to be developed. This ranking algorithm is called R-RCA. From Equations 21 and 22, it can be observed that the calculations of the inverse of the n × n matrix and the multiplication between two n × n matrices in the algorithm 1 This can be successfully avoided. However, the parameter space is doubled. The relaxation effectively improves computing performance.

[0093] A softmax normalization according to an embodiment of the present invention is now described.

[0094] Here, the product of the rank order values ​​of two nodes i and j, r is used. i ▪r j, is used as strength of evidence that the edge between these two nodes vanishes (is broken). However, it suffers from the extreme values ​​or outliers in the rank order values ​​r. To reduce the influence of extreme values ​​or outliers in the data without removing them from the dataset, softmax normalization is applied to the rank order values ​​r. Thus, the rank order values ​​are nonlinearly transformed using the S-function before the two rank order values ​​are multiplied. Therefore, the recovery error is ‖ ( σ ( r ) σ ( r ) T ) ∘ M − P ˜ ‖ F 2 and σ(▪) is the softmax function with the following: σ ( r ) i = e r i ∑ k = 1 n e r k , ( i = 1, … n ) .

[0095] The corresponding objective function for the algorithm 1 will be changed to the following: min e ≥ 0 ‖ ( σ ( Be ) σ T ( Be ) ∘ M − P ˜ ) ‖ F 2 + τ ‖ e ‖ 1 .

[0096] Similarly, the objective function for equation 20 is changed to the following: min e ≥ 0, r ≥ 0 cr T ( I n − A ˜ ) r + ( 1 − c ) ‖ r − e ‖ F 2 + λ ‖ ( σ ( r ) σ T ( r ) ) ∘ M- P ˜ ‖ F 2 + τ ‖ e ‖ 1 .

[0097] The optimization of these two objective functions is based on the following two theorems.

[0098] Sentence 4 The update of e according to equation 26 as follows reduces the objective function in equation 24 monotonically until convergence: e ← e ∘ { 4 [ ( B T Ψ P ˜ ) ∘ M ] σ Be 4 [ ( B T Ψσ ( Be ) σ ( Be ) ) ∘ M ] σ ( Be ) + τ 1 n } 1 4 , where Ψ ={diag [σ(Be)] - σ(Be)σ T (Be)} is.

[0099] Sentence 5 The update of r according to equation 27 as follows reduces the objective function in equation 25 monotonically until convergence: r ← r ∘ { c 2 A ˜ r + λ [ ( ( σ ( r ) 1 n T ) ∘ P ˜ + ρΛ ) ∘ M ] σ ( r ) + ( 1 − c ) 2 e 1 2 r + λ [ ( ( σ ( r ) ∘ σ ( r ) σ T ( r ) + σ ( r ) ( σ T ( r ) P ˜ ) ) ∘ M ) σ ( r ) ] } 1 4 , with Λ = σ(r) σ T (r) and ρ = σ T (r)σ(r).

[0100] The sentence 4 and the sentence 5 can use a similar strategy to the set 1 This can be proven. The ranking algorithms with softmax normalization are referred to as RCA-SOFT and R-RCA-SOFT, respectively.

[0101] A description of smoothing by temporally and dynamically broken grids according to an embodiment of the present invention is now given.

[0102] Over time, anomalies can propagate within the invariant network. However, the underlying causal anomalies typically remain unchanged for a given time period T. Based on this intuition, a smoothing procedure is developed by jointly considering temporally and dynamically broken networks. That is, a smoothing term is added to the previously described objective functions. ‖ e ( i ) − e ( i-1 ) ‖ 2 2 added. Here are e (i-1) and e (i) The rank order vectors of causal anomalies at two continuation times. The objective function of the RCA algorithm for smoothing time-broken networks is then shown in Equation 28 as follows: min e ( i ) ≥ 0,1 ≤ i ≤ T ∑ i = 1 T [ ‖ ( Be ( i ) ( e ( i ) ) T B T ) ∘ M − P ˜ ‖ F 2 + τ ‖ e ( i ) ‖ 1 ] + α ‖ e ( i ) − e ( i − 1 ) ‖ 2 2 .

[0103] The update formula of equation 28 can be derived as follows: e ( i ) ← e ( i ) ∘ { 4 [ ( B T P ˜ ) ∘ M ] Be+2 α e ( i − 1 ) 4 [ ( B T Bee T B T ) ∘ M ] Be + τ 1 n +2 α e ( i ) } 1 4 .

[0104] The ranking algorithms with time-domain smoothing are designated as T-RCA, TR-RCA, T-RCA-SOFT and TR-RCA-SOFT, respectively.

[0105] A description of the features / advantages of the present invention compared to conventional methods will now be given.

[0106] Existing approaches for detecting causal anomalies using invariant networks share three common limitations: ( 1 ) They do not take into account the potential propagation of disturbances in the invariant network; ( 2 ) the ranking assessment guideline they used is not good evidence of a causal anomaly; and ( 3They cannot jointly consider the temporally and dynamically broken networks that are recognized as useful for noise reduction. The inventors' approach overcomes these limitations by explicitly considering the propagation of the initial seed anomalies using the newly restarted random walk model, and by enforcing that the anomalies are temporally smooth across adjacent timestamps.

[0107] A description of the competing / commercial value of the solution achieved by the present invention will now be given.

[0108] The present invention can improve the accuracy of detecting true causal anomalies in large systems such as power plants, cloud computing, production lines, computer network systems, etc. By identifying the most critical anomaly, human operators can save considerable effort in testing, maintenance, and repair of large physical systems. This can increase the uptime of large systems and thus their output.

[0109] The embodiments described here can be entirely hardware, entirely software, or contain both hardware and software elements. In a preferred embodiment, the present invention is implemented in software, which includes, but is not limited to, firmware, resident software, microcode, etc.

[0110] Embodiments may include a computer program product accessible from a computer-usable or computer-readable medium that provides program code for use by or in conjunction with a computer or any other instruction execution system. A computer-usable or computer-readable medium may include any device that stores, transmits, propagates, or transports the program for use by or in conjunction with an instruction execution system, instruction execution device, or instruction execution apparatus. The medium may be a magnetic, optical, electronic, electromagnetic system, an infrared or semiconductor system (or a magnetic, optical, electronic, electromagnetic device or apparatus, an infrared or semiconductor device or apparatus), or a propagation medium.The medium can contain a computer-readable storage medium such as semiconductor or solid-state memory, magnetic tape, removable computer disk, read / write memory (RAM), read-only memory (ROM), magnetic hard disk, optical disk, etc.

[0111] Each computer program can be specifically stored in machine-readable storage media or in a machine-readable storage device (e.g., a program memory or a magnetic disk) that is readable by a programmable general-purpose or specialized computer to configure and control the operation of a computer in order to execute the procedures described herein when the storage media or storage device is read by the computer. Furthermore, the system according to the invention can be considered to be embodied in a computer-readable storage medium configured with a computer program, wherein the storage medium is configured to cause a computer to operate in a specific and predetermined manner to perform the functions described herein.

[0112] A data processing system capable of storing and / or executing program code may contain at least one processor directly or indirectly coupled to memory elements via a system bus. The memory elements may include local memory used during the actual execution of the program code, mass storage, and cache memory, which provides temporary storage of at least some program code to reduce the number of times code is read from mass storage during execution. Input / output or I / O devices (including, but not limited to, keyboards, displays, pointers, etc.) may be coupled to the system either directly or via intermediate I / O controllers.

[0113] Furthermore, network adapters can be connected to the system to allow the data processing system to be connected to other data processing systems or remote printers or storage devices via private or public intermediate networks. Modems, cable modems, and Ethernet cards are just some of the currently available types of network adapters.

[0114] References to "exactly one embodiment" or "an embodiment" of the present invention, as well as variants thereof, in the description mean that a specific feature, structure, property, etc., described in connection with the embodiment, is included in at least one embodiment of the present invention. Therefore, occurrences of the expression "in exactly one embodiment" or "in an embodiment," as well as any other variants appearing at various points throughout the patent specification, do not necessarily all refer to the same embodiment.

[0115] It will be appreciated that the use of any of the following: " / ", "and / or", and "at least one of", e.g., in cases of "A / B", "A and / or B", and "at least one of A and B", should include the selection of only the first executed option (A), or the selection of only the second executed option (B), or the selection of both options (A and B). As a further example, in cases of "A, B and / or C" and "at least one of A, B and C", this wording should include the selection of only the first listed option (A), or the selection of only the second listed option (B), or the selection of only the third listed option (C), or the selection of only the first and second listed options (A and B), or only the selection of the first and third listed options (A and C), or only the selection of the second and third listed options (B and C), or the selection of all three options (A and B and C).As the average professional in this and related fields can easily understand, this can be extended to many of the listed positions.

[0116] Naturally, the foregoing is intended to be illustrative and exemplary in every respect, but not limiting, and the scope of protection of the invention disclosed herein is not to be determined from the detailed description, but instead from the claims interpreted in accordance with the full breadth permitted by patent law. Naturally, the embodiments shown and described herein are merely illustrative of the principles of the present invention, and a person skilled in the art can implement various modifications without deviating from the scope of protection or the inventive concept of the invention. A person skilled in the art could implement various other combinations of features without deviating from the scope of protection or the inventive concept of the invention.Having thus described aspects of the invention in detail and with the comprehensiveness required by patent law, the attached claims set out what is claimed by the patent specification and for which protection is sought. QUOTES INCLUDED IN THE DESCRIPTION

[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited patent literature

[0000] US 62 / 292383

[0001]

Claims

[1] Computer-implemented method for root cause anomaly detection in an invariant network with multiple nodes generating time series data, wherein the method comprises: Modeling anomaly propagation in the invariant network using a processor; Reconstructing broken invariant links in an invariant graph based on rank-order vectors of causal anomalies by the processor, wherein each of the broken invariant links comprises a respective pair of nodes formed from the multiple nodes such that one of the nodes in the respective pair of nodes has an anomaly, wherein each of the rank-order vectors of causal anomalies serves to specify, for a given of the multiple nodes when arranged in pairs, a respective node anomaly status; The processor calculates a sparse penalty of the rank order vectors of causal anomalies to obtain a set of time-dependent anomaly orders; Performing a time-smoothing operation on the set of time-dependent anomaly orders by the processor; and Controlling an anomaly initiation of one of the multiple nodes based on the set of time-dependent anomaly orders by the processor. [2] Computer-implemented method according to claim 1, wherein the anomaly propagation in the invariant network is modeled using a random walk restart technique. [3] Computer-implemented method according to claim 1, wherein the anomaly propagation in the invariant network is modeled on the basis of a threshold applied to a triple consisting of a first degree of an AutoRegressive-eXogenous model (AXR model), a second degree of the AXR model and a time delay between the time series data. [4] Computer-implemented method according to claim 3, further comprising: Comparing possible combinations of values ​​of the triples for a given pair of observed values ​​of the time series data for a given pair of nodes formed from the multiple nodes, in order to determine respective suitability ratings for each of several ARX models formed for the given pair of nodes; and Selecting one of the several ARX models with the highest of the respective suitability ratings. [5] Computer-implemented method according to claim 3, further comprising: Calculating a suitability score for the ARX model for a given pair of observed values ​​of the time series data for a given pair of nodes formed from the multiple nodes; and Identifying an invariant link for the given pair of nodes as broken based on the suitability assessment for the ARX model for the given pair of observed values ​​that exceeds a suitability threshold. [6] Computer-implemented method according to claim 1, wherein the sparse penalty of the rank order vectors of causal anomalies is used to control a number of zero distinct values ​​in the set of time-dependent anomaly orders. [7] Computer-implemented method according to claim 1, further comprising determining a recovery error for the invariant network based on the broken invariant connections, wherein the set of time-dependent anomaly arrangements is determined based on the recovery error. [8] Computer-implemented method according to claim 1, further comprising identifying the broken invariant connections based on an objective function. [9] Computer-implemented method according to claim 1, wherein the anomaly propagation in the invariant network is modeled using an objective function based on an anomaly weight vector after propagation of an initial disturbance in the rank order vectors of causal anomalies. [10] Computer-implemented method according to claim 1, wherein the control step comprises switching off a computer processing device in the case of an anomaly initiating multiple nodes in order to mitigate error propagation thereof. [11] Computer-implemented method according to claim 1, wherein the control step comprises selectively terminating a root cause process, which is executed in a computer processing device in which an anomaly initiating multiple nodes is performed, in order to mitigate error propagation thereof. [12] Computer-implemented method according to claim 1, further comprising optimizing the set of time-dependent anomaly arrangements. [13] Computer-implemented method according to claim 12, wherein the set of time-dependent anomaly orderings is optimized using an objective function with a first term requiring consistency between a propagated anomaly and a broken invariant combination, and with a second term being a penalty term to promote zero entries in the rank order vectors of causal anomalies. [14] Computer program product for root cause anomaly detection in an invariant network with multiple nodes generating time series data, wherein the computer program product comprises a non-transitory computer-readable storage medium embodying program instructions, wherein the program instructions are executable by a computer to cause the computer to execute a procedure comprising: Modeling anomaly propagation in the invariant network using a processor; Reconstructing broken invariant links in an invariant graph based on rank-order vectors of causal anomalies by the processor, wherein each of the broken invariant links comprises a respective pair of nodes formed from the multiple nodes such that one of the nodes in the respective pair of nodes has an anomaly, wherein each of the rank-order vectors of causal anomalies serves to specify, for a given of the multiple nodes when arranged in pairs, a respective node anomaly status; The processor calculates a sparse penalty of the rank order vectors of causal anomalies to obtain a set of time-dependent anomaly orders; Performing a time-smoothing operation on the set of time-dependent anomaly orders by the processor; and Controlling an anomaly initiation of one of the multiple nodes based on the set of time-dependent anomaly orders by the processor. [15] Computer program product according to claim 14, wherein the anomaly propagation in the invariant network is modeled on the basis of a threshold applied to a triple formed from a first degree of an AutoRegressive-eXogenous model (AXR model), a second degree of the AXR model and a time delay between the time series data. [16] Computer program product according to claim 15, wherein the method further comprises: Calculating a suitability score for the ARX model for a given pair of observed values ​​of the time series data for a given pair of nodes formed from the multiple nodes; and Identifying an invariant link for the given pair of nodes as broken based on the suitability assessment for the ARX model for the given pair of observed values ​​that exceeds a suitability threshold. [17] Computer program product according to claim 14, wherein the sparse penalty of the rank order vectors of causal anomalies is used to control a number of zero distinct values ​​in the set of time-dependent anomaly orders. [18] Computer program product according to claim 14, wherein the method further comprises determining a recovery error for the invariant network based on the broken invariant connections, wherein the set of time-dependent anomaly arrangements is determined based on the recovery error. [19] Computer program product according to claim 14, wherein the anomaly propagation in the invariant network is modeled using an objective function based on an anomaly weight vector after propagation of an initial disturbance in the rank order vectors of causal anomalies. [20] Computer processing system for root cause anomaly detection in an invariant network with multiple nodes generating time series data, wherein the system comprises: a processor configured to: Modeling anomaly propagation in the invariant network; Reconstructing broken invariant connections in an invariant graph based on rank-order vectors of causal anomalies, wherein each of the broken invariant connections comprises a respective pair of nodes formed from the multiple nodes such that one of the nodes in the respective pair of nodes has an anomaly, wherein each of the rank-order vectors of causal anomalies serves to specify, for a given of the multiple nodes when arranged in pairs, a respective node anomaly status; Computing a sparse penalty of the rank order vectors of causal anomalies to obtain a set of time-dependent anomaly orders; Performing a time-smoothing operation on the set of time-dependent anomaly orders; and Controlling the initiation of an anomaly at one of the multiple nodes based on the set of time-dependent anomaly orders.

Citation Information

Patent Citations

  • Ranking Causal Anomalies via Temporal and Dynamical Analysis on Vanishing Correlations

    US62292383P0

  • Method and Apparatus for Performing Capacity Planning and Resource Optimization in a Distributed System

    US20080228459A1

  • Fault Localization in Distributed Systems Using Invariant Relationships

    US20140047279A1

  • Anomaly detection in spatial and temporal memory system

    US20140067734A1