Industrial Internet of Things fault root cause positioning method based on causal diagram
Through a causal graph-based method, combined with mRMR, Lasso regression and random walk algorithm, the problem that traditional methods are difficult to accurately identify the root cause of failure in complex systems is solved, and efficient and accurate fault root cause positioning and path display are achieved.
Patent Information
- Application Number
- CN202510595864.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-08-01
AI Technical Summary
The existing traditional fault diagnosis methods are difficult to accurately identify the root cause of the fault in complex systems, especially in the absence of system topology information. Traditional causal analysis algorithms cannot effectively capture time dependence, resulting in reduced performance of the root cause positioning task.
A causal graph-based method is adopted, combining the variable screening method of mRMR and Lasso regression, and a causal graph is generated using PC algorithm and partial cross-mapping, and a root cause of failure is located through a random walk algorithm, and a root cause of failure is determined by combining the state transition probability matrix and the number of visits sorting.
It improves the accuracy and efficiency of fault root cause positioning, generates a more accurate and complete causal map, provides clear fault propagation paths and dependencies, and enhances the robustness and adaptability of the system.
Smart Images

Figure CN120416013A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer technology and relates to a method for locating the root cause of industrial Internet of Things faults based on causal graphs. Background Art
[0002] In industrial processes centered around complex interconnected systems and components, accurately identifying and promptly correcting process faults is of utmost importance. If these faults are not resolved promptly, they may lead to operational disruptions, increased safety risks, and economic losses. Although traditional fault diagnosis methods are still effective in many cases, they often struggle to trace the deep - seated root causes when faced with certain complex faults. Merely detecting a fault is not sufficient to completely solve the problem; delving into its root cause is the key. Given the limitations of traditional diagnosis methods, the demand for innovative root - cause diagnosis technologies in modern industrial manufacturing is more urgent than ever. Most existing research focuses on fault detection techniques, but the more challenging root - cause diagnosis (i.e., being able to identify causal relationships between variables, trace fault propagation paths, and precisely locate the root cause) still faces many problems. Traditional root - cause diagnosis methods usually rely on experts' experience and knowledge, making it difficult to develop accurate root - cause diagnosis models. With the continuous progress of sensor technology and data acquisition means, a large amount of process data can be efficiently recorded and stored, which provides new research opportunities for data - based root - cause diagnosis methods and makes them gradually become an important direction in the current field. The construction and analysis of causal graphs have become important tools for fault tracing. These graphs can reveal the dependencies in the time series after a fault occurs, helping analysts clarify the causes and propagation paths of faults, thereby providing more targeted diagnostic support.
[0003] Since many complex systems can only provide raw KPI data rather than system topologies, finding the dependencies between components by constructing a fault impact graph is a very important part of the root - cause analysis system. The construction of the fault impact graph mainly uses causal analysis algorithms. However, the causal analysis algorithms used in existing root - cause analysis research usually assume independent and identically distributed data, such as the PC algorithm. This assumption makes it difficult for the algorithm to capture the temporal dependence of time - series data. This approach results in the inability to consider the influence of temporal dependence when inferring the causal relationships of time - series data, and further leads to the causal graph being unable to accurately represent the dependencies between components, causing a reduction in the performance of the root - cause location task. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a method for locating the root cause of faults in industrial Internet of Things based on causal graphs, to solve the problem of automatically identifying the root cause of faults in industrial Internet of Things. At the same time, a variable screening method combining mRMR and Lasso regression is used, and combined with the PC algorithm and partial cross mapping to generate causal graphs. The random walk algorithm combining causality and correlation is used to effectively improve the accuracy of root cause location.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A method for locating the root cause of faults in industrial Internet of Things based on causal graphs, the method comprising the following steps:
[0007] S1. When a fault occurs in the industrial Internet of Things, extract the time series data of the key performance indicators of the industrial Internet of Things;
[0008] S2. Calculate the redundancy degree of variables using the maximum correlation and minimum redundancy criterion to reduce the feature redundancy; then extract the key features through the least absolute shrinkage and selection operator regression;
[0009] S3. Construct a causal graph establishment model combining the traditional PC algorithm and partial cross mapping to generate a causal graph between the index data;
[0010] S4. Establish a state transition probability matrix based on correlation and causality;
[0011] S5. Walk on the causal graph according to the state transition probability matrix and record the access times; according to the access times, sort and finally determine the root cause of the fault.
[0012] Further, in step S1, the key performance indicator data related to the industrial Internet of Things is collected in real time. The key performance indicator data in the industrial Internet of Things includes CPU-related data, disk-related data, memory-related data, and network-related data.
[0013] Further, in step S2, it specifically includes the following steps:
[0014] S21. Filter the redundancy of the obtained key performance indicator data through the maximum correlation and minimum redundancy method, and screen out the key variable set X that has a positive effect on root cause analysis r ;
[0015] S22. Then use the least absolute shrinkage and selection operator method to further screen and optimize the selected key features to obtain the final feature set representation.
[0016] Further, in step S21, it is set that X and Y are the original data, then the maximum dependence relationship between the two is expressed as:
[0017]
[0018] Among them, I(·, *) is the mutual information function, and x i represents the i-th variable in X, and m represents the number of variables included;
[0019] The minimum redundancy criterion for both to select mutually exclusive variables is expressed as:
[0020]
[0021] Among them, x i , x j respectively represent the i-th and j-th variables in X;
[0022] Then, the mutual information difference and the mutual information quotient are used to obtain the variables related to the anomaly, where the mutual information difference is expressed as:
[0023]
[0024] Among them, r is the number of anomaly-related variables obtained,
[0025] The mutual information quotient is expressed as:
[0026]
[0027] Thus, the set of anomaly-related variables obtained through the mutual information difference maxφ1(x, y) and the mutual information quotient maxφ2(x, y) is denoted as X r .
[0028] Furthermore, in step S22, the L1 criterion is introduced as a penalty term on the regression coefficient and added to the sum of squares of the residuals,
[0029] which is specifically expressed as:
[0030]
[0031] Among them, y is the dependent variable; X r is the set of variables after mRMR screening; β is the regression coefficient to be evaluated, λ is the parameter controlling the regularization strength, and ||·||1 represents the L1 norm. The set of features after screening and optimization by the lasso method is denoted as X’ r .
[0032] Furthermore, in step S3, the process of constructing the causal graph includes:
[0033] S31. Given a set of observed variables X1, X2,..., X n , construct the wireless graph skeleton through the PC algorithm. Initially, assume that all variable pairs (X i , X j) Each has an edge, generating a complete undirected graph;
[0034] S32. By checking whether each pair of variables (X i , X j ) is conditionally independent, gradually delete the insignificant edges, and finally obtain the undirected graph skeleton E;
[0035] S33. Orient the relationship between each pair of variables through the partial cross mapping method, and determine whether there is a causal relationship between the two according to whether the high-order cross mapping error between variables is higher than the threshold Q th ;
[0036] S34. Combine the PC algorithm and the partial cross mapping orientation process to obtain a complete directed causal graph G=(V, E), where V represents the set of all variables, and E represents the directed causal relationship between variables.
[0037] Furthermore, in step S33, during the process of deleting edges, start from each node and conduct a conditional independence test on other nodes; if each pair of variables (X i , X j ) is conditionally independent given a certain set, then delete the edge between them; by continuously conducting conditional independence tests and deleting edges, finally obtain an undirected graph skeleton; the process is expressed as:
[0038]
[0039] In the formula, indicates that two variables are conditionally independent, where represents the conditional independence relationship; S represents the given conditional set.
[0040] Furthermore, in step S34, for each pair of variables (X, Y) to be analyzed for relationship orientation, apply the partial cross mapping method to it and other variables different from this pair of variables to detect the direct causal relationship between the two, and exclude the possible indirect causal confounding relationship outside the two variables. Among them, the high-order cross mapping error of the two variable values is calculated as follows:
[0041]
[0042] Among them, represents other variables except variables X and Y; m is the number of additional variables, represents the cross mapping value of X and Y, represents the cross mapping value under the influence of the remaining variables excluded; after obtaining this value through the partial cross mapping method, compare it with the threshold Q th to determine. If it is higher than the threshold, there is a causal relationship, otherwise there is no causal relationship.
[0043] Further, in step S4, the following steps are included:
[0044] S41. Initialize the state transition matrix Q;
[0045] S42. Design three walking directions, including forward transfer, reverse transfer, and staying at the origin. Among them, the probability of each walking direction is determined according to the DTW correlation coefficient between two adjacent nodes; among them,
[0046] The process of calculating the DTW correlation coefficient between two adjacent nodes in the causal graph is:
[0047] Corr DTW = DTW(x i , x j )
[0048] In the formula, x i , x j represent two adjacent nodes, and DTW(·) is the DTW function;
[0049] Then the forward transfer probability is set as:
[0050] p ij = Corr DTW + αQ ij
[0051] Among them, α is the weight factor for controlling causality; Q ij is the causal coefficient from node i to node j;
[0052] The reverse transfer probability is set as:
[0053] p ji = ρ(Corr DTW + αQ ij )
[0054] Among them, ρ ∈ [0, 1) is the discount factor, and the value of it indicates whether to allow the walker to visit each node in the weighted fault influence graph;
[0055] The design of staying at the origin is: when the similarity between the current node where the walker is located and the abnormal node is higher than that of other neighbor nodes, the staying time of the walker at the current node is determined by the difference between the maximum probability of its upstream variable and the maximum probability of its downstream variable:
[0056]
[0057] In the formula, e kj represents the edge between k and j, and p ij represents the transfer probability from i to j;
[0058] Finally, normalize each row of the matrix to obtain the final state transition matrix.
[0059] Further, in step S5, perform a random walk on the causal graph. The walker in the random walk algorithm iteratively walks on the causal graph multiple times through the transition probability, and counts the visit times to obtain the visit count array C[m], where c i ∈C[m], i = 1,..., m represents the visit count;
[0060] In the process of sorting according to the visit count and finally determining the root cause of the fault, the larger the c i of a certain node, the greater the probability that it is the root cause; sort the visit counts in descending order, and list the top Top-k root causes of the fault as the root cause identification result.
[0061] The beneficial effects of the present invention are as follows:
[0062] The present invention first uses the mRMR algorithm to preliminarily screen the original data, effectively removing redundant and irrelevant features, and retaining the key variables highly related to the root cause of the fault. Subsequently, the variables are further refined and optimized through Lasso regression, ensuring the representativeness and discriminant ability of the selected variables. This process greatly improves the efficiency and accuracy of subsequent causal graph generation.
[0063] In the causal graph generation stage, the present invention adopts the PC algorithm and the partial cross-mapping technique. The PC algorithm infers the causal relationship between variables through conditional independence tests and constructs a preliminary undirected graph. The partial cross-mapping technique further orients and refines the undirected graph to generate a more accurate and complete causal graph. This causal graph not only intuitively shows the dependence relationship between each node, but also provides a clear path for locating the root cause of the fault.
[0064] The present invention also introduces a random walk algorithm that combines causality and correlation. Based on the causal graph, this algorithm dynamically analyzes and adjusts the dependence relationship between nodes through random walks, making the dependence relationship in the causal graph more accurately and comprehensively expressed. This innovative algorithm not only improves the accuracy of root cause location of faults, but also enhances the robustness and adaptability of the causal graph.
[0065] Other advantages, objectives and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following specification. Brief Description of the Drawings
[0066] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail and preferably below in conjunction with the accompanying drawings, where:
[0067] Figure 1 is the overall flowchart of the industrial Internet of Things fault root cause location method based on the causal graph according to the embodiment of the present invention;
[0068] Figure 2 is the framework schematic diagram of the industrial Internet of Things fault root cause location method based on the causal graph according to the embodiment of the present invention;
[0069] Figure 3 is the schematic diagram of the causal graph generation process according to the embodiment of the present invention. Specific Embodiments
[0070] The following uses specific specific examples to illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0071] Among them, the drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and should not be construed as a limitation to the present invention; in order to better illustrate the embodiments of the present invention, some components in the drawings will be omitted, enlarged, or reduced, which does not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0072] In the drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or position relationship, they are based on the orientation or position relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the position relationship in the drawings are only for illustrative purposes and should not be construed as a limitation to the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0073] Please refer to Figures 1 to 3 , which is an industrial Internet of Things fault root cause location method based on the causal graph.
[0074] The present invention provides a method for fault root cause location based on causal graph in the industrial Internet of Things scenario. Please refer to Figure 1 , the method includes:
[0075] S1. When a fault occurs in the industrial Internet of Things, extract the time series data of the key performance indicators of the industrial Internet of Things;
[0076] S2. Calculate the redundancy degree of variables using the maximum correlation and minimum redundancy criterion to reduce the feature redundancy; then extract the key features through the least absolute shrinkage and selection operator regression;
[0077] S3. Construct a causal graph by combining the traditional PC algorithm and partial cross mapping to establish a model, and generate the causal graph between each index data;
[0078] S4. Obtain the state transition probability matrix based on the correlation and causality;
[0079] S5. Walk on the causal graph and record the access times; sort according to the access times and finally determine the fault root cause.
[0080] In a preferred embodiment, step S1 includes: real-time collecting the key performance indicator (KPI) data related to the industrial Internet of Things. Specifically, the KPI data in the industrial Internet of Things includes CPU-related data, disk-related data, memory-related data, and network-related data. In this embodiment, different fault types are simulated by means of fault injection to generate abnormal behaviors, so as to obtain the key performance indicator data containing fault information.
[0081] In a preferred embodiment, step S2 includes:
[0082] S21. First, use the maximum correlation and minimum redundancy method mRMR to filter the redundancy of the obtained key performance indicator data, efficiently filter out the redundant variables, and accurately screen out the key variables that are of great value to the root cause analysis; assume that X and Y are the original data, then the maximum dependence relationship can be defined as:
[0083]
[0084] where I(·,*) is the mutual information function, and x i represents the i-th variable in X, and m represents the number of variables included.
[0085] The minimum redundancy criterion for selecting mutually exclusive variables is:
[0086]
[0087] where x i , x j represent the i-th and j-th variables in X respectively.
[0088] Use the mutual information difference and the mutual information quotient to obtain variables related to anomalies. The formula for the mutual information difference is as follows:
[0089]
[0090] where r is the number of variables related to the obtained anomalies.
[0091] The formula for the mutual information quotient is as follows:
[0092]
[0093] Thus, the set of variables related to anomalies obtained through the mutual information difference maxφ1(x,y) and the mutual information quotient maxφ2(x,y) is denoted as X r .
[0094] S22. For the fault subset screened by the mRMR method, use the least absolute shrinkage and selection operator method lasso for more refined feature screening and optimization. Introduce the L1 criterion as a penalty term into the sum of squared residuals of the regression coefficients to increase the predictive ability of the ordinary least squares estimation. Its form is:
[0095]
[0096] where y is the dependent variable; X r is the set of variables after mRMR screening; β is the regression coefficient to be evaluated, λ is the parameter controlling the regularization strength, and ||·||1 represents the L1 norm. When using the L1 criterion, the irrelevant coefficients in β will be compressed to 0, and the corresponding variables will be screened out, while the variables corresponding to the non-zero coefficients are considered candidate fault variables. The feature set after screening and optimization by the lasso method is denoted as X r '.
[0097] In a preferred embodiment, step S3 includes: constructing a causal graph by combining the traditional PC algorithm (Peter-Clark, PC) and partial cross mapping, as shown in the causal graph construction process of the combination of PC and partial cross mapping method in Figure 3 The process is as follows:
[0098] S31. Use the PC algorithm to construct an undirected graph skeleton. Given a set of observed variables X1, X2,..., X n , initially, assume that there are edges for all variable pairs (X i ,X j ) in the graph, and generate a complete undirected graph.
[0099] S32. By testing each pair of variables (X i ,X jWhether it is conditionally independent, gradually delete the insignificant edges. Specifically, starting from each node, perform a conditional independence test on other nodes. If X i and X j are conditionally independent given a certain set, then delete the edge between them. By continuously performing conditional independence tests and deleting edges, an undirected graph skeleton is finally obtained. In this graph, the edges between variables indicate a certain degree of dependence between them, and the conditional independence test excludes irrelevant edges. Finally, the undirected graph skeleton is obtained:
[0100]
[0101] wherein, indicates that two variables are conditionally independent, where represents the conditional independence relationship; S represents the given conditional set.
[0102] S33. After the undirected graph skeleton is constructed, the relationship between each pair of variables is oriented by the partial cross-mapping method. For each pair of variables (X, Y) to be analyzed, the partial cross-mapping method is applied to different variables except this pair of variables to detect the direct causal relationship between them and exclude the possible indirect causal confounding relationship existing outside the two variables. The high-order cross-mapping error of the values of the two variables is calculated as follows:
[0103]
[0104] wherein, represents other variables except variables X and Y; m is the number of additional variables, represents the cross-mapping value of X and Y, represents the cross-mapping value under the condition of excluding the influence of the remaining variables. After obtaining this value through the partial cross-mapping method, it is compared with the threshold Q th to determine. If it is higher than the threshold, there is a causal relationship, otherwise there is no causal relationship.
[0105] S34. Finally, combining the PC algorithm and the partial cross-mapping orientation process, a complete directed causal graph G=(V, E) is obtained, where V represents the set of all variables, and E represents the directed causal relationship between variables. This causal graph reveals the causal interaction structure between variables in the system and can reflect the propagation path when an anomaly occurs.
[0106] In a preferred embodiment, step S4 includes:
[0107] S41. First, initialize the state transition matrix Q.
[0108] S42. To prevent the random walk algorithm from getting trapped in a low-abnormality area during the walk and being unable to escape, three walking directions are designed: forward transfer, reverse transfer, and staying at the origin;
[0109] First, calculate the DTW correlation coefficient between two adjacent nodes in the causal graph. The calculation formula is as follows:
[0110] Corr DTW = DTW(x i , x j )
[0111] In the formula, x i , x j represent two adjacent nodes, and DTW(·) is the DTW function.
[0112] The probability calculations for the three walking directions are as follows:
[0113] (1) Forward transfer: The forward transfer probability is set as:
[0114] p ij = Corr DTW + αQ ij
[0115] where α is a weight factor for controlling causality; Q ij is the causality coefficient from node i to node j.
[0116] (2) Reverse transfer: The reverse transfer probability is set as:
[0117] p ji = ρ(Corr DTW + αQ ij )
[0118] where ρ ∈ [0, 1) is a discount factor. If the value of ρ is small, it restricts the walker from walking on the path of the weighted fault influence graph; if the value of ρ is large, it means allowing the walker to visit various nodes in the weighted fault influence graph;
[0119] (3) Staying at the origin: When the similarity between the current node where the walker is located and the abnormal node is higher than that of other neighbor nodes, the time for the walker to stay at the current node is determined by the difference between the maximum probability of its upstream variable and the maximum probability of its downstream variable:
[0120]
[0121] In the formula, e kj represents the edge between k and j, and p ij represents the transfer probability from i to j.
[0122] Finally, normalize each row of the matrix to obtain the final state transition matrix.
[0123] In a preferred embodiment, step S5 includes: performing a walk on the causal graph. The walker in the random walk algorithm iteratively walks on the causal graph multiple times through the transition probability, and counts the number of visits to obtain an array of visit counts C[m], where c i ∈C[m], i = 1,..., m represents the number of visits.
[0124] In the process of sorting according to the number of visits and finally determining the root cause of the fault, the larger the c i of a certain node, the greater the probability that it is the root cause. Sort the number of visits in descending order, and list the top Top-k fault root causes as the result of root cause identification.
[0125] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A method for locating the root cause of faults in industrial Internet of Things based on causal diagrams, characterized in that: The method includes the following steps: S1. When a failure occurs in the industrial Internet of Things, extract the time series data of the key performance indicators of the industrial Internet of Things; S2. Calculate the redundancy degree of variables using the maximum correlation and minimum redundancy criterion to reduce the feature redundancy; then extract the key features through the least absolute shrinkage and selection operator regression; S3. Construct a causal graph by combining the PC algorithm and the partial cross mapping to establish a model, and generate a causal graph between the data of each index; S4. Establish a state transition probability matrix based on the correlation and causality; S5. Walk on the causal graph according to the state transition probability matrix, record the access times; and finally determine the root cause of the failure according to the access times.
2. The industrial Internet of Things fault root cause location method based on a causal graph according to claim 1, wherein: In step S1, the key performance indicator data related to the industrial Internet of Things is collected in real time. The key performance indicator data in the industrial Internet of Things includes CPU-related data, disk-related data, memory-related data, and network-related data.
3. A method for fault root cause location in industrial Internet of Things based on causal graph according to claim 1, characterized in that: In step S2, it specifically includes the following steps: S21. Redundancy filtering is performed on the obtained key performance indicator data by the maximum correlation and minimum redundancy method to screen out a set of key variables X that have a positive effect on root cause analysis. r ; S22. Then use the least absolute shrinkage and selection operator method to further screen and optimize the selected key features to obtain the final feature set representation.
4. The industrial Internet of Things fault root cause location method based on a causal graph according to claim 3, characterized in that: In step S21, it is set that X and Y are the original data, then the maximum dependence relationship between the two is expressed as: where I(·,*) is the mutual information function, and x i represents the i-th variable in X, and m represents the number of variables included; The minimum redundancy criterion for selecting mutually exclusive variables between the two is expressed as: where x i , x j represent the i-th and j-th variables in X, respectively; Then use the mutual information difference and mutual information quotient to obtain the variables related to the anomaly, where the mutual information difference is expressed as: where r is the number of variables related to the obtained anomaly, The mutual information quotient is expressed as: Thus, the set of variables related to anomalies obtained through the mutual information difference maxφ1(x,y) and the mutual information quotient maxφ2(x,y) is denoted as X r .
5. The industrial Internet of Things fault root cause location method based on a causal diagram according to claim 3, characterized in that: In step S22, an L1 criterion is introduced into the regression coefficient as a penalty term and added to the sum of the squares of the residuals, and its specific expression is: Among them, y is the dependent variable; X r is the set of variables after mRMR screening; β is the regression coefficient to be evaluated, λ is the parameter controlling the regularization strength, ||·||1 represents the L1 norm; the feature set after screening and optimization by the lasso method is denoted as X’ r .
6. The industrial Internet of Things fault root cause location method based on a causal graph according to claim 3, wherein: In step S3, the process of constructing the causal graph includes: S31. Given a set of observed variables X1, X2, …, X n , construct a wireless graph skeleton through the PC algorithm. Initially, assume that there is an edge for all variable pairs (X i , X j ) in the graph to generate a complete undirected graph; S32. By checking whether each pair of variables (X i , X j ) is conditionally independent, gradually delete the insignificant edges, and finally obtain an undirected graph skeleton E; S33. Orient the relationship between each pair of variables by means of a partial cross-mapping method, and determine whether there is a causal relationship between the two according to whether the high-order cross-mapping error between the variables is higher than a threshold Q th ; S34. Combine the PC algorithm and the partial cross mapping orientation process to obtain a complete directed causal graph G=(V, E), where V represents the set of all variables, and E represents the directed causal relationship between variables.
7. A root cause location method for industrial Internet of Things faults based on a causal graph according to claim 6, characterized in that: In step S33, during the process of deleting edges, starting from each node, perform a conditional independence test on other nodes; if each pair of variables (X i , X j ) is conditionally independent given a certain set, then delete the edge between them; by continuously performing conditional independence tests and deleting edges, finally obtain an undirected graph skeleton; the process is expressed as: In the formula, indicates that two variables are conditionally independent, where represents the conditional independence relationship; S represents the given set of conditions.
8. A method for locating the root cause of faults in an industrial Internet of Things based on a causal diagram according to claim 7, characterized in that: In step S34, for each pair of variables (X, Y) for which relationship orientation analysis is to be performed, the partial cross-mapping method is used to detect the direct causal relationship between it and variables different from this pair of variables, and the possible indirect causal confounding relationship existing outside the two variables is excluded, where the high-order cross-mapping error of the values of the two variables is calculated as shown below: Among them, represents other variables except variables X and Y; m is the number of additional variables, represents the cross-mapping value of X and Y, represents the cross-mapping value under the condition of excluding the influence of the remaining variables; after obtaining this value through the partial cross-mapping method, it is compared with the threshold Q th to determine. If it is higher than the threshold, there is a causal relationship; otherwise, there is no causal relationship.
9. The industrial Internet of Things fault root cause location method based on a causal diagram according to claim 6, wherein: In step S4, it includes the following steps: S41. Initialize the state transition matrix Q; S42. Design three walking directions, including forward transfer, reverse transfer, and staying at the origin. Among them, the probability of each walking direction is determined according to the DTW correlation coefficient between two adjacent nodes; among them, The process of calculating the DTW correlation coefficient between two adjacent nodes in the causal graph is: Corr DTW = DTW(x i , x j ) where x i , x j represent two adjacent nodes, and DTW(·) is the DTW function; Then the forward transfer probability is set as: p ij = Corr DTW + αQ ij Among them, α is the weight factor for controlling causality; Q ij is the causal coefficient from node i to node j; The reverse transfer probability is set as: p ji = ρ(Corr DTW + αQ ij ) where ρ∈[0, 1) is a discount factor, and its value indicates whether to allow the walker to visit each node in the weighted fault impact graph; The origin staying is designed as: when the similarity between the current node where the walker is located and the abnormal node is higher than that of other neighbor nodes, the staying time of the walker at the current node is determined by the difference between the maximum probability of its upstream variable and the maximum probability of its downstream variable: where e kj represents the edge between k and j, and p ij represents the transition probability from i to j; Finally, normalize each row of the matrix to obtain the final state transition matrix.
10. A method for fault root cause location in industrial Internet of Things based on causal diagram according to claim 9, characterized in that: In step S5, a random walk is performed on the causal graph. The walker in the random walk algorithm iteratively walks on the causal graph multiple times through the transition probability, and counts the number of visits to obtain the array C[m] of visit counts, where c i ∈C[m], i = 1, ..., m represents the number of visits; In the process of sorting according to the number of accesses and finally determining the root cause of the failure, the larger the c of a certain node i is, the greater the probability that it is the root cause; sort the number of accesses in descending order and list the top Top-k root causes of the failure as the root cause identification result.
Citation Information
Cited By
Industrial chain risk monitoring method, system and equipment based on causal analysis and medium
CN121352519A