Fault root cause analysis method and apparatus for multiple systems, and computer device
By constructing a multi-system network and dynamically adjusting the correlation weights, the problem of ignoring the relationship between devices in traditional detection methods is solved, enabling accurate location of the root cause of the fault and improvement of production efficiency.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-29
- Publication Date
- 2026-03-05
AI Technical Summary
Traditional anomaly detection methods in steel production ignore the interdependencies between equipment, leading to ineffective inspections and delays in fault identification, which affects production efficiency and product quality.
A multi-system network is constructed. By receiving data and updating the correlation weights of the edges, the time delay parameters are calculated using a regression model. Combined with Bayes' theorem and forgetting factor, the correlation is dynamically adjusted to construct a Bayesian network for fault root cause analysis.
Accurately pinpoint the root cause of a fault, reduce unnecessary inspections, improve the efficiency of fault location and repair, reduce operating costs, and adapt to changes in the production environment.
Smart Images

Figure CN2024115531_05032026_PF_FP_ABST
Abstract
Description
Multi-system root cause analysis methods, devices, and computer equipment Technical Field
[0001] This application relates to the field of fault analysis, and more specifically, to a method, apparatus, computer equipment, and storage medium for multi-system fault root cause analysis. Background Technology
[0002] In steel production, large pieces of equipment such as motors, drive systems, frames, and rolls form the core of production activities, interconnected to support complex and precise operations. Traditional anomaly detection methods typically focus on monitoring and analyzing individual measurement points on each piece of equipment. Once the data at a measurement point deviates from the normal range, an alarm is triggered, prompting workers to conduct on-site inspections of the corresponding equipment. However, this point-to-point anomaly detection strategy ignores the interdependencies between equipment; that is, an anomaly in a piece of equipment parameter may not be a signal of a fault in that equipment itself, but rather a result of a fault in related equipment. Neglecting these indirect connections can lead to ineffective inspections, wasted human resources, delays in identifying and repairing the true fault, and impacts production efficiency and product quality.
[0003] Summary of the Invention
[0004] This summary section is provided to introduce some selected concepts in a simplified form, which will be further described in the detailed description section below. This summary section is not intended to identify any key or essential features of the claimed subject matter, nor is it intended to help determine the scope of the claimed subject matter.
[0005] Based on this, this application discloses a multi-system fault root cause analysis method, which includes:
[0006] Constructing nodes and edges of a multi-system network; wherein the nodes represent measurement points for anomaly detection in the multi-system, and the edges represent the correlations between the nodes;
[0007] Receive data from multiple systems, and update the weights of the edges with respect to the correlation by calculating the data;
[0008] Based on the nodes and edges of the multi-system network, list the nodes sorted according to their relevance.
[0009] By using the above methods, the root cause of the fault can be accurately located, which can not only reduce unnecessary inspections, but also speed up the fault location and repair process, ultimately promoting increased production efficiency and reduced operating costs.
[0010] Furthermore, it also includes: calculating the time delay parameters between the nodes through a regression model; wherein the time delay parameters represent the one-way causal relationship, the two-way causal relationship, or the independent relationship between the nodes.
[0011] By using the above method, the time delay parameter can be used to understand the temporal relationship between multiple nodes, which helps to identify the correlation between nodes and facilitates the analysis of the root cause of the fault.
[0012] Furthermore, receiving multi-system data and updating the weights of the edges with respect to the correlation by calculating the data includes:
[0013] Receive first-time data or partial data, and update the weight of the edge with respect to the correlation based on the first-time data or partial data.
[0014] By using the above method and updating the edge weights with the latest or partial data, the weights of correlations can be precisely adjusted, which can help with subsequent root cause analysis of faults.
[0015] Furthermore, receiving first-time data or partial data, and updating the weights of the edges with respect to the correlation using the first-time data or partial data, includes:
[0016] By incorporating a forgetting factor, the weighting of the first time-phase data or a portion of the data on the correlation is adjusted.
[0017] By adjusting the impact of newly received data on the original weights in the above manner, the specific correlation between the nodes can be adjusted more accurately, which can help with subsequent root cause analysis of faults.
[0018] Furthermore, receiving multi-system data and updating the weights of the edges with respect to the correlation by calculating the data includes:
[0019] Receive multi-system data, and according to Bayes' theorem, use the multi-system data as posterior probability data to update the weights of the edges with respect to the correlation.
[0020] The above method allows for more convenient and accurate adjustment of the edge weights, thereby more accurately adjusting the correlation values and aiding in subsequent root cause analysis of faults.
[0021] Furthermore, after receiving data from multiple systems, the process includes cleaning and preprocessing the data from the multiple systems.
[0022] By using the above methods, data can be processed to obtain data that conforms to the corresponding rules, which is helpful for subsequent correlation calculations and improves effectiveness.
[0023] Furthermore, this application also discloses a multi-system fault root cause analysis device, which includes:
[0024] A construction module is used to construct nodes and edges of a multi-system network; wherein the nodes represent measurement points for anomaly detection in the multi-system network, and the edges represent the correlations between the nodes.
[0025] An update module is used to receive data from multiple systems and update the weights of the edges with respect to the correlation by calculating the data;
[0026] The display module is used to list the nodes in a sorted manner based on the nodes and edges of the multi-system network.
[0027] This application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described method.
[0028] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method.
[0029] This application also provides a computer program product tangibly stored on a computer-readable medium and comprising computer-executable instructions that, when executed, cause at least one processor to perform the methods described above. Attached Figure Description
[0030] Implementations of this disclosure are illustrated in the accompanying drawings by way of example rather than limitation, and similar reference numerals in the drawings denote the same or similar parts.
[0031] Figure 1 is a schematic flowchart of a method for multi-system fault root cause analysis according to an embodiment of this application.
[0032] Figure 2 is a schematic diagram of a multi-system fault root cause analysis apparatus according to an embodiment of this application.
[0033] Figure 3 is a schematic diagram of a computer device for multi-system fault root cause analysis according to an embodiment of this application.
[0034] Figure 4 is a schematic diagram of a directed graph for fault root cause analysis of a multi-system according to an embodiment of this application.
[0035] The reference numerals in the accompanying drawings are as follows: S101-S103 Step 200: Device 201: Module 202: Module 203: Module 204: Module 300: Computer device 302: Processor 304: Memory Detailed Implementation
[0036] In the following description, numerous specific details are set forth for illustrative purposes. However, it will be understood that the invention can be implemented without these specific details. In other examples, well-known circuits, structures, and techniques have not been shown in detail so as not to affect the understanding of the description.
[0037] Throughout the specification, references to "an implementation," "implementation," "exemplary implementation," "some implementations," "various implementations," etc., indicate that the implementation of the invention described may include specific features, structures, or characteristics. However, it is not implied that every implementation must include these specific features, structures, or characteristics. Furthermore, some implementations may have some, all, or none of the features described for other implementations.
[0038] The systematic root cause analysis algorithm proposed in this application promotes a shift from reactive to proactive prevention by deeply understanding the interactions between equipment, thus helping steel companies develop smarter and more efficient equipment management systems. By accurately locating the root cause of faults, unnecessary inspections can be reduced, and the fault location and repair process can be accelerated, ultimately leading to increased production efficiency and reduced operating costs. Furthermore, the algorithm's dynamic update capability allows it to flexibly adapt to changes in the production environment, providing a new perspective and practical solutions for equipment maintenance and management in the steel industry and potentially other heavy industries.
[0039] Specifically, this application discloses a method for root cause analysis of faults in multiple systems, comprising:
[0040] S101, Construct nodes and edges of a multi-system network; wherein the nodes represent measurement points for anomaly detection in the multi-system, and the edges represent the correlations between the nodes.
[0041] The multiple systems mentioned above can be the steel production system, composed of multiple pieces of equipment, such as motors, drive systems, frames, etc. Furthermore, based on the study of the steel production line process and operations, relevant measurement points are connected by lines to construct a preliminary directed graph. The direction of the arrows represents the cause-and-effect relationship. As shown in Figure 4, nodes 1, 2, 3, 4, 5, and 6 represent different measurement points, and the directed lines between nodes represent different associations between them. The magnitude of the correlation between nodes will be further indicated later.
[0042] Specifically, constructing a measurement point correlation model can be achieved by using historical data and industry expert rules to calculate the correlation between various equipment measurement points, forming a directed graph model where nodes represent equipment measurement points and edge weights reflect the strength of the correlation between measurement points.
[0043] S102, Receive data from multiple systems, and update the weight of the edge with respect to the correlation by calculating the data.
[0044] Specifically, receiving data from multiple systems and updating the weights of the edges with respect to the correlation by calculating the data includes:
[0045] Receive first-time data or partial data, and update the weight of the edge with respect to the correlation based on the first-time data or partial data.
[0046] The first time-based data can represent the latest data generated at the aforementioned measurement points. Partial data can represent data generated only from a subset of measurement points used for correlation calculation. By using the latest or partial data to update the edge weights, the correlation weights can be precisely adjusted, thus aiding subsequent root cause analysis. In the above steps, receiving the first time-based data or partial data, and updating the edge weights with respect to the correlation using the first time-based data or partial data, includes: adjusting the influence of the first time-based data or partial data on the correlation weights by adding a forgetting factor. By adjusting the influence of newly received data on the original weights, the specific correlations between the nodes can be further accurately adjusted, aiding subsequent root cause analysis. Furthermore, after receiving multi-system data, the process includes: cleaning and preprocessing the multi-system data to obtain data conforming to relevant rules, which helps in subsequent correlation calculations and improves effectiveness.
[0047] In some implementations, correlation analysis focuses on each target measurement point for anomaly detection and performs joint temporal analysis on the complete measurement point data that is theoretically correlated to it, in order to find the measurement points with the strongest correlation. The degree of correlation between two measurement points x and y can be calculated by calculating the contribution of their respective data changes to each other, mathematically represented as...
[0048] Furthermore, the steps can also be described as a dynamic incremental updating mechanism, that is, throughout the entire process of the algorithm's operation, this application uses a dynamic incremental method to continuously update the directed graph model to ensure that it reflects the real-time changes in the device state and improves the accuracy and timeliness of anomaly detection.
[0049] Furthermore, it also includes: calculating the time delay parameters between the nodes through a regression model; wherein the time delay parameters represent the one-way causal relationship, the two-way causal relationship, or the independent relationship between the nodes.
[0050] Specifically, the time-series inference part determines the time lag parameters between the relevant data of multiple paired measurement points. Then, by establishing an autoregressive model and comparing the prediction model errors under different measurement point data, dynamic causal inference based on time series is achieved, helping to discover the root cause of anomalies. Taking two time series variables x and y as an example, joint autoregressive analysis is performed on them. At any given time, the lag parameters of x and y are represented as α, β, γ, δ, respectively, and the constant terms are represented as u_{1t} and u_{2t}. Therefore, we further obtain:
[0051] Based on the model's learning results, there are four possible causal scenarios:
[0052] (1) X is the cause of the change in Y, indicating a one-way causal relationship from X to Y. If the estimated coefficient of lagged x in equation (1) is statistically significant, while the estimated coefficient of lagged y in equation (2) is not statistically significant, then x is considered the cause of the change in y.
[0053] (2) y is the cause of the change in x, indicating a one-way causal relationship from y to x. If the estimated coefficient of lagged y in equation (2) is statistically significant, while the estimated coefficient of lagged x in equation (1) is not statistically significant, then y is considered the cause of the change in x.
[0054] (3) X and Y have a bidirectional causal relationship, which means a unidirectional causal relationship from X to Y and a unidirectional causal relationship from Y to X. If the estimated coefficient of lagged x in equation (1) is statistically significant as a whole, and the estimated coefficient of lagged y in equation (2) is also statistically significant as a whole, then it is considered that there is a feedback relationship or a bidirectional causal relationship between x and y.
[0055] (4) X and Y are independent, or there is no causal relationship between X and Y. If the estimated coefficient of lagged x in equation (1) is not significant as a whole, and the estimated coefficient of lagged y in equation (2) is also not significant as a whole, then it is considered that there is no causal relationship between x and y.
[0056] By using the above method, the time delay parameter can be used to understand the temporal relationship between multiple nodes, which helps to identify the correlation between nodes and facilitates the analysis of the root cause of the fault.
[0057] In some implementations, receiving multi-system data and updating the weights of the edges with respect to the correlation by calculating the data includes: receiving the multi-system data and, according to Bayes' theorem, using the multi-system data as posterior probability data to update the weights of the edges with respect to the correlation. This method allows for more convenient and accurate adjustment of the edge weights, thereby more accurately adjusting the correlation values and aiding in subsequent root cause analysis of faults.
[0058] Based on the statistical analysis above, nodes and directed line segments can be added to complete a preliminary directed graph. The direction of the lines is from cause to effect.
[0059] In summary, based on the baseline of the measurement point, the baseline deviation trend of the measurement point is monitored in real time, and an alarm is issued when abnormal points occur frequently. The data-based system analysis can be summarized into three parts: correlation analysis, temporal inference, and graph model update. In some implementations, the graph model update part includes constructing a Bayesian network and continuously updating it; the network update uses an incremental learning process to simplify computation.
[0060] S103, Based on the nodes and edges of the multi-system network, list the nodes sorted by correlation.
[0061] Specifically, the above steps can be called Systematized Root Cause Analysis: This algorithm, based on a constructed directed graph model, identifies multiple potential causes of anomalies at specific measurement points and sorts them by probability, providing clear guidance for on-site inspections by field personnel and significantly improving the efficiency and accuracy of fault diagnosis.
[0062] By using the above methods, the root cause of the fault can be accurately located, which can not only reduce unnecessary inspections, but also speed up the fault location and repair process, ultimately promoting increased production efficiency and reduced operating costs.
[0063] Furthermore, this application primarily introduces a root cause analysis algorithm specifically tailored for multi-equipment systems in the steel industry. This algorithm aims to reveal potential causal relationships between equipment by combining data-driven methods and rule-based inputs, thereby accurately pinpointing the causes behind anomalies. In contrast, some existing technologies rely on industry expertise and Granger causality analysis to construct a Bayesian network (directed acyclic graph), supporting incremental updates. Previous industrial root cause analysis studies have employed methods such as Failure Mode and Effects Analysis (FMEA), Fault Tree Analysis (FTA), and Reliability Block Diagrams (RBD) to construct Bayesian networks. These methods heavily depend on business logic and cannot accurately quantify various factors as the method presented in this application. Furthermore, the Bayesian network in this application can dynamically adapt and update according to changes in operating conditions and processes, which is the innovation of the method in this application.
[0064] In some implementations, the overall network update algorithm of this application flows as follows:
[0065] 1. Initialization:
[0066] 1.1 Define the network structure: Determine the nodes (variables) and edges (conditional dependencies) in the network.
[0067] 1.2 Initialization parameters: Assign initial values to each conditional probability distribution. These can be uniform distributions or prior probabilities based on domain knowledge.
[0068] 2. Receive new data:
[0069] 2.1 Data Preprocessing: Clean and preprocess the newly received data to ensure that they meet the model input requirements.
[0070] 2.2 Instantiation: Instantiate new data as the state of observed variables in the network.
[0071] 3. Update parameters:
[0072] 3.1 Calculate the probability: For each conditional probability table, calculate the probability under the corresponding condition based on the new data.
[0073] 3.2 Application Update Rules:
[0074] 3.2.1 Online Expectation-Maximization (EM) algorithm: If the EM algorithm is used for parameter estimation, an online version can be adopted, which uses only the first time data point or a small batch of data, instead of all historical data, to update sufficient statistical data.
[0075] 3.2.2 Incremental Bayesian Update: The posterior probability is updated directly using Bayes' formula, that is, the new data is treated as the observation and the conditional probability table is updated.
[0076] 3.2.3 Forgetting Factor: Introducing a forgetting factor α (0 < α ≤ 1), the update combines old probabilities with new evidence.
[0077] New probability = α * probability of new evidence + (1-α) * old probability, to balance the influence of new and old data.
[0078] 4. Performance Evaluation:
[0079] 4.1 Periodic validation: Use cross-validation or discard some data as a validation set to evaluate the model's performance on new data.
[0080] 4.2 Monitoring metrics: Observe performance metrics such as accuracy and AUC-ROC (Area Under Curve-receiver operating characteristic) to ensure that the model does not overfit or its performance degrades.
[0081] 5. Decision-making and adjustments:
[0082] 5.1 Structure Adjustment: If a decline in model performance is observed, it may be necessary to consider adjusting the network structure. This is usually quite complex and may require additional algorithmic support.
[0083] 5.2 Parameter Adjustment: Fine-tune hyperparameters, such as learning rate and forgetting factor, based on the validation results.
[0084] 6. Loop:
[0085] Repeat the above steps until the stopping condition is met, such as reaching a predetermined number of learning rounds, a performance improvement threshold, or the end of the data flow.
[0086] The multi-system fault root cause analysis method disclosed in this application has the following beneficial effects:
[0087] 1. Construction of Bayesian networks with precise quantification
[0088] The method proposed in this application leverages industry expertise combined with Granger causality analysis to construct Bayesian networks, offering significant advantages over traditional root cause analysis techniques such as FMEA, FTA, and RBD. Traditional methods heavily rely on business logic when constructing Bayesian networks, often failing to accurately quantify the relationships between factors. Our method, however, combines deep industry insights with Granger causality analysis. This integration ensures more precise quantification of causal relationships within the network, thereby improving the accuracy and reliability of root cause analysis.
[0089] 2. Dynamic adaptive update mechanism
[0090] Another key innovation lies in the dynamic adaptive updating capability of the Bayesian network described in this application. In industrial environments, operating conditions and processes frequently change, placing higher demands on the real-time performance and effectiveness of root cause analysis. Unlike past static analysis models, the Bayesian network proposed in this paper can automatically update and optimize according to actual changes in operating conditions and processes. This dynamic characteristic ensures that the network model always reflects the latest system state, thereby improving the efficiency of fault diagnosis and prevention, and providing strong support for the maintenance and optimization of industrial systems.
[0091] These two innovations together constitute the core advantage of the method presented in this application, not only improving the accuracy of root cause analysis but also enhancing its adaptability and practicality. They provide a new perspective and toolkit for problem-solving in the industrial field, marking a significant advancement in this area.
[0092] To determine whether the methods described in this application have been used, the following can be tested: First, if another algorithm claims to employ a "causal reasoning" method, particularly Granger causality analysis, for diagnosing equipment failures in the steel industry, the probability of suspicion increases by 20%. Second, if another algorithm claims its "graph" or "Bayesian network" can be automatically updated without requiring historical data input, the probability of suspicion also increases by 20%. Third, if the heavy machinery covered by the other algorithm includes not only mechanical equipment but also main drive systems such as thyristors or insulated-gate bipolar transistors (IGBTs), the probability of suspicion surges by 50%.
[0093] In some embodiments, steel industry plants have gradually completed initial digitalization in recent years, accumulating a certain amount of data. Therefore, plants have begun to introduce anomaly prediction and root cause analysis technologies. However, in the field of large equipment anomaly detection and root cause analysis, existing technology and equipment providers primarily focus on individual devices, rather than performing anomaly prediction and root cause analysis on multiple interconnected devices. However, many customers desire joint analysis of their equipment. Furthermore, the physical connections between multiple devices within the same plant make joint analysis more accurate in predicting equipment problems and identifying the root causes behind them.
[0094] Based on market research, it was found that customers have a large demand for system anomaly detection that takes into account device relationships. Therefore, this application proposes an architectural framework for device system anomaly detection, which also supports the integration of artificial intelligence solutions and personalized business rules.
[0095] The methods and apparatus mentioned in this application have three key features: 1. They focus on predictive maintenance and root cause analysis of steel industry facilities; 2. They employ statistical methods and deterministic rules for causal reasoning; and 3. They consider the interrelationships between multiple measurement points on different facilities.
[0096] The entire algorithm consists of four main parts:
[0097] 1. Training the anomaly detection model
[0098] Algorithm 1: Training the prediction model
[0099] - Collect and preprocess historical data of motors and transmission equipment.
[0100] - Align data by time period.
[0101] - The prediction model is trained using the Gaussian Mixture Model (GMM) algorithm.
[0102] - Save the trained prediction model.
[0103] 2. Anomaly detection and risk warning:
[0104] Algorithm 2: Prediction and Alarm
[0105] - Collect and preprocess new data from motors and transmission equipment.
[0106] - Align the new data by time period.
[0107] - Use a trained GMM prediction model to predict the confidence interval for new data.
[0108] - Compare the collected data with the predicted data.
[0109] - A risk alert is triggered when the amount of abnormal data reaches a threshold.
[0110] 3. Establish a causal reasoning model:
[0111] Algorithm 3: Establishing a causal reasoning model
[0112] - Collect and preprocess historical data of anomalies.
[0113] - Use the Granger causal model to determine the correlation of the data.
[0114] - Establish a basic directed graph model and use Bayesian network methods to complete the weights.
[0115] - Save the graph model.
[0116] 4. Root Cause Analysis
[0117] Algorithm 4: Root Cause Analysis
[0118] - Collect and preprocess new anomalous data.
[0119] - Use the previously established graphical model to perform root cause analysis.
[0120] In some embodiments, parts 3 and 4 constitute the main body of the multi-system fault root cause analysis method described above in this application, while parts 1 and 2 can be preliminary preparations for establishing a directed graph model and collecting data from measurement points.
[0121] For example, measurement points for a motor may include temperature, current, bearing pressure, load rate, and motor vibration. Measurement points for a transmission facility may include temperature, load current, load voltage, and thyristor temperature. Based on historical data for these measurement points, a directed graph is constructed using Granger Causality (GC) and a Bayesian network. Then, according to the aforementioned correlation calculations, edge weights are assigned, and finally, the nodes are sorted by correlation to identify the nodes with the highest correlation. Each node (measurement point) corresponds to one or more root causes of a fault, such as abnormal power supply voltage, abnormal power supply current, motor overload, motor load imbalance, and risk of bearing failure. In some embodiments, once the most relevant node (measurement point) is identified, the underlying root cause of the fault becomes clear.
[0122] It should be understood that although the steps in the flowchart of Figure 1 are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in Figure 1 may include multiple steps or multiple stages, which are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps.
[0123] Figure 2 provides an apparatus 200 for multi-system fault root cause analysis. The apparatus 200 includes:
[0124] The construction module 201 is used to construct nodes and edges of a multi-system network; wherein the nodes represent measurement points for anomaly detection in the multi-system, and the edges represent the correlations between the nodes.
[0125] Update module 202 is used to receive multi-system data and update the weight of the edge with respect to the correlation by calculating the data;
[0126] The display module 203 is used to list the nodes sorted by correlation based on the nodes and edges of the multi-system network.
[0127] It should be noted that the device may contain more or fewer modules to implement the described functions. For example, at least one module in FIG2 may be further divided into a plurality of different sub-modules, each sub-module being used to perform at least a portion of the operations described herein in conjunction with the corresponding module. Furthermore, in some examples, device 200 may also include additional modules for performing other operations already described in the specification. Moreover, those skilled in the art will understand that the exemplary device 200 may be implemented using software, hardware, firmware, or any combination thereof.
[0128] Figure 3 provides a computer device. According to one embodiment, the computer device 300 may include a processor 302 that executes a computer program stored in a memory 304. When executed by the processor, the computer program implements the method described above.
[0129] Those skilled in the art will understand that the structure shown in Figure 3 is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or may combine certain components, or may have different component arrangements.
[0130] Those skilled in the art will understand that all or part of the processes in the methods described above can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0131] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the above steps.
[0132] This application also provides a computer program product tangibly stored on a computer-readable medium and including computer-executable instructions that, when executed, cause at least one processor to perform the methods described above.
[0133] Furthermore, the computer program can be stored and run in the cloud to execute the method. Furthermore, the components of the program can be deployed on multiple devices or in the cloud; for example, corresponding steps can be deployed and run on a local computer, or run on different cloud devices, transmitting signals via communication connections, or they can also be deployed and run on a local computer. This application does not limit the described approach or method; corresponding technologies can be flexibly deployed and fully utilized to execute and complete the method using cloud computing, big data, supercomputing capabilities, and other equipment and technologies.
[0134] Some implementations of this disclosure may include an article of writing. The article of writing may include a storage medium for storing logic. Examples of storage media may include one or more types of computer-readable storage media capable of storing electronic data, including volatile or non-volatile memory, removable or non-removable memory, erasable or non-erasable memory, writable or rewritable memory, and so on. Examples of logic may include various software units, such as software components, programs, applications, computer programs, application programs, system programs, machine programs, operating system software, middleware, firmware, software modules, routines, subroutines, functions, methods, procedures, software interfaces, application programming interfaces (APIs), instruction sets, computational code, computer code, code segments, computer code segments, words, values, symbols, or any combination thereof. In some implementations, for example, the article of writing may store executable computer program instructions that, when executed by a processor, cause the processor to perform the methods and / or operations described herein. Executable computer program instructions may include any suitable type of code, such as source code, compiled code, interpreted code, executable code, static code, dynamic code, and so on. Executable computer program instructions can be implemented according to a predefined computer language, method, or syntax used to command the computer to perform specific functions. These instructions can be implemented using any suitable high-level, low-level, object-oriented, visual, compiled, and / or interpreted programming language.
[0135] The examples described above include those of the disclosed architecture. It is certainly impossible to describe every conceivable combination of components and / or methods, but those skilled in the art will understand that many other combinations and arrangements are also possible. Therefore, this novel architecture is intended to cover all such alternatives, modifications, and variations that fall within the spirit and scope of the appended claims.
Claims
1. A method for root cause analysis of faults in multiple systems, wherein, include: Constructing nodes and edges for multi-system networks; Wherein the nodes represent measurement points for anomaly detection in the multi-system, and the edges represent the correlation between the nodes; Receive data from multiple systems, and update the weights of the edges with respect to the correlation by calculating the data; Based on the nodes and edges of the multi-system network, list the nodes in the order of their correlations.
2. The method according to claim 1, wherein, Also includes: The time delay parameters between the nodes are calculated using a regression model; wherein the time delay parameters represent the one-way causal relationship, the two-way causal relationship, or the independent relationship between the nodes.
3. The method according to claim 1, wherein, Receiving data from multiple systems and updating the weights of the edges with respect to the correlation by calculating the data includes: Receive first-time data or partial data, and update the weight of the edge with respect to the correlation based on the first-time data or partial data.
4. The method according to claim 3, wherein, Receiving first-time data or partial data, and updating the weights of the edges with respect to the correlation based on the first-time data or partial data, including: By incorporating a forgetting factor, the weighting of the first time-phase data or a portion of the data on the correlation is adjusted.
5. The method according to claim 1, wherein, Receiving data from multiple systems and updating the weights of the edges with respect to the correlation by calculating the data includes: Receive multi-system data, and according to Bayes' theorem, use the multi-system data as posterior probability data to update the weights of the edges with respect to the correlation.
6. The method according to claim 1, wherein, After receiving data from multiple systems, the following is included: The data from the multiple systems are cleaned and preprocessed.
7. A multi-system fault root cause analysis device (200), wherein, include: Module (201) is used to build nodes and edges in a multi-system network; The nodes represent the multiple systems. Measurement points for anomaly detection, where the edges represent the correlation between the nodes; The update module (202) is used to receive multi-system data and update the weight of the edge with respect to the correlation by calculating the data; The display module (203) is used to list the nodes in the order of the correlation based on the nodes and edges of the multi-system network.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
10. A computer program product tangibly stored on a computer-readable medium and comprising computer-executable instructions that, when executed, cause at least one processor to perform the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Fault detection method and system for servo inverter
CN118245917A
Power grid fault power protection maintenance plan making method, device, equipment and medium
CN118504939A
Interdependent causal networks for root cause localization
US20230069074A1
Systematic prognostic analysis with dynamic causal model
WO2020046261A1
Fault diagnosis method and system based on qualitative trend analysis and five-state bayesian network, and device and storage medium
WO2023071220A1