Alarm root cause positioning method and device, equipment, storage medium and program product

By performing cluster analysis and frequency analysis on alarm data sets in a cloud environment, establishing a fault propagation diagram, and using the spanning tree principle to locate the root cause, the problem of insufficient accuracy in alarm root cause location in existing technologies is solved, and automated and accurate root cause location is achieved.

CN120611203APending Publication Date: 2025-09-09CHINA MOBILE INFORMATION TECHNOLOGY CO LTD +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510743125.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing technologies lack accuracy in locating the root causes of multiple alarms in complex businesses in cloud environments, making it difficult to effectively address the challenges of human subjectivity, rule limitations, and complexity analysis.

Method used

By performing cluster analysis on the alarm data set, clustering algorithms such as FCM are used to cluster the alarm data, combining frequency analysis with the establishment of a fault propagation graph, and using the spanning tree principle to locate the root cause, the root cause of the alarm data set is determined.

Benefits of technology

It reduces the difficulty of root cause location, realizes automatic and accurate root cause location, and improves the efficiency and accuracy of locating multiple alarms for complex services in cloud environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120611203A_ABST
    Figure CN120611203A_ABST
Patent Text Reader

Abstract

The invention provides an alarm root cause positioning method and device, equipment, a storage medium and a program product, and relates to the technical field of intelligent operation and maintenance. The method comprises the following steps: performing clustering analysis on an alarm data set to obtain a plurality of alarm data subsets; performing frequency analysis on each alarm data subset to obtain a frequency analysis result of each alarm data subset; according to the multiple frequency analysis results, a fault propagation graph corresponding to the alarm data set is established, the fault propagation graph comprises multiple nodes and edges connecting the two nodes, the nodes represent the alarm data, and the edges represent the causal relationship between the two alarm data; and performing root cause positioning on the fault propagation graph by adopting a spanning tree principle, and determining a root cause corresponding to the alarm data set. According to the scheme, automatic and accurate root cause positioning can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent operation and maintenance technology, and specifically to an alarm root cause locating method, device, equipment, storage medium and program product. Background Art

[0002] Existing methods for locating the root cause of alarms often employ rule-based troubleshooting and rules based on alarm knowledge graphs. Rule-based troubleshooting uses predefined rules or rule bases to determine the relationship between alarms and root causes. These rules can be based on specific failure modes, system configurations, or indicator thresholds. Operations and maintenance personnel then match and judge the rules to identify the likely root cause. While this approach can help identify the root cause of alarms to a certain extent, it is limited by human subjectivity, rule limitations, and the challenges of complex analysis, making it difficult to accurately locate the root cause of multiple alarms in complex business environments. Summary of the Invention

[0003] At least one embodiment of the present application provides a method, apparatus, device, storage medium, and program product for locating an alarm root cause, which are used to solve the problem in the prior art that it is difficult to accurately locate the root cause of an alarm root cause.

[0004] In order to solve the above technical problems, this application is implemented as follows:

[0005] In a first aspect, an embodiment of the present application provides a method for locating an alarm root cause, including:

[0006] Perform cluster analysis on the alarm data set to obtain multiple alarm data subsets;

[0007] Performing frequency analysis on each of the alarm data subsets to obtain a frequency analysis result for each of the alarm data subsets;

[0008] Establishing a fault propagation graph corresponding to the alarm data set based on the plurality of frequency analysis results, the fault propagation graph comprising a plurality of nodes and edges connecting two of the nodes, wherein the nodes represent the alarm data and the edges represent a causal relationship between the two alarm data;

[0009] The root cause of the fault propagation graph is located by using a spanning tree principle to determine the root cause corresponding to the alarm data set.

[0010] Optionally, the alarm root cause locating method, wherein frequency analysis is performed on each of the alarm data subsets to obtain a frequency analysis result corresponding to each of the alarm data subsets, includes:

[0011] Using a Gaussian function as a kernel function to obtain a probability density function corresponding to each of the alarm data subsets;

[0012] According to the probability density function, obtaining a first alarm data subset having a probability density greater than a preset probability density in each of the alarm data subsets;

[0013] Perform frequency analysis on the first alarm data subset to obtain a frequency analysis result of each of the alarm data subsets.

[0014] Optionally, the alarm root cause location method, wherein establishing a fault propagation graph corresponding to the alarm data set based on the plurality of frequency analysis results, includes:

[0015] According to the plurality of frequency analysis results, obtaining a plurality of first alarm data in the alarm data set, the plurality of first alarm data having a number greater than a preset number or a frequency greater than a preset frequency;

[0016] Perform correlation analysis on the plurality of first alarm data sets to establish a fault propagation graph corresponding to the alarm data set.

[0017] Optionally, the alarm root cause location method, wherein the root cause location is performed on the fault propagation graph using a spanning tree principle to determine the root cause corresponding to the alarm data set, includes:

[0018] If there is a closed loop in the fault propagation graph, the pruning rule in the spanning tree principle is used to prune the loop and locate the root cause of the fault propagation graph to determine the root cause corresponding to the alarm data set.

[0019] Optionally, the alarm root cause location method, wherein cluster analysis is performed on the alarm data set to obtain multiple alarm data subsets, includes:

[0020] Randomly selecting a plurality of second alarm data from the alarm data set as initial cluster centers;

[0021] For each of the alarm data except the plurality of second alarm data in the alarm data set, obtaining a degree of membership of the alarm data relative to each of the initial cluster centers;

[0022] Iteratively updating the initial cluster center according to the membership degree to obtain multiple cluster centers after iteration;

[0023] A plurality of alarm data sets are obtained based on the plurality of cluster centers.

[0024] Optionally, in the alarm root cause location method, the alarm data set is obtained by preprocessing, and the preprocessing includes at least one of the following:

[0025] Data cleaning; format conversion; feature extraction; data standardization; data normalization.

[0026] In a second aspect, an embodiment of the present application provides an alarm root cause location device, including:

[0027] A first acquisition module is used to perform cluster analysis on the alarm data set to obtain multiple alarm data subsets;

[0028] a second obtaining module, configured to perform frequency analysis on each of the alarm data subsets to obtain a frequency analysis result of each of the alarm data subsets;

[0029] An establishment module is used to establish a fault propagation graph corresponding to the alarm data set based on the plurality of frequency analysis results, wherein the fault propagation graph includes a plurality of nodes and an edge connecting two of the nodes, wherein the nodes represent the alarm data and the edges represent a causal relationship between the two alarm data;

[0030] The determination module is used to locate the root cause of the fault propagation graph using a spanning tree principle to determine the root cause corresponding to the alarm data set.

[0031] In a third aspect, an embodiment of the present application provides an alarm root cause locating device, comprising: a processor, a memory, and a program or instruction stored on the memory and executable on the processor, wherein when the processor executes the program or instruction, the alarm root cause locating method as described in the first aspect is implemented.

[0032] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method for locating the root cause of the alarm as described in the first aspect is implemented.

[0033] In a fifth aspect, an embodiment of the present application provides a computer program product, comprising computer instructions, which, when executed by a processor, implement the alarm root cause locating method as described in the first aspect.

[0034] Compared with the prior art, the embodiments of the present application provide a method, device, storage medium and program product for locating the root cause of an alarm, the method comprising: performing cluster analysis on an alarm data set to obtain multiple subsets of alarm data; performing frequency analysis on each of the alarm data subsets to obtain frequency analysis results for each of the alarm data subsets; establishing a fault propagation graph corresponding to the alarm data set based on the multiple frequency analysis results, the fault propagation graph comprising multiple nodes and edges connecting two of the nodes, wherein the nodes represent the alarm data and the edges represent the causal relationship between the two alarm data; performing root cause location on the fault propagation graph using the spanning tree principle to determine the root cause corresponding to the alarm data set. In this way, by performing root cause location through cluster analysis, frequency analysis and fault propagation graph, the difficulty of root cause location is reduced, and automated and accurate root cause location is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present application. The same reference symbols are used throughout the drawings to represent the same components. In the drawings:

[0036] Figure 1 Schematic diagram of the process of locating the root cause of an alarm in an embodiment of the present application;

[0037] Figure 2 Schematic diagram of the process flow of the FCM algorithm in the embodiment of the present application;

[0038] Figure 3 Schematic diagram of the FCM algorithm membership in an embodiment of the present application;

[0039] Figure 4 A schematic diagram of the center point of the FCM algorithm in an embodiment of the present application;

[0040] Figure 5 This is a schematic diagram of the FCM algorithm running data in the embodiment of the present application;

[0041] Figure 6 Schematic diagram of the FCM algorithm operation results in the embodiment of the present application;

[0042] Figure 7 A schematic diagram of a fault propagation diagram in an embodiment of the present application;

[0043] Figure 8 This is a flow chart of one implementation of the alarm root cause location method in an embodiment of the present application;

[0044] Figure 9 Schematic diagram of another embodiment of the alarm root cause location method in the embodiment of the present application;

[0045] Figure 10 Schematic diagram of the structure of the alarm root cause locating device in an embodiment of the present application;

[0046] Figure 11 This is a hardware block diagram of the alarm root cause locating device in an embodiment of the present application. DETAILED DESCRIPTION

[0047] The terms "first", "second", etc. in this application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the terms used in this way are interchangeable where appropriate, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are generally of the same type, and do not limit the number of objects, for example, the first object can be one or more. In addition, "or" in this application represents at least one of the connected objects. For example, "A or B" covers three options, namely, Option 1: including A but not including B; Option 2: including B but not including A; Option 3: including both A and B. The character " / " generally indicates that the objects associated before and after are in an "or" relationship.

[0048] Please refer to Figure 1 , an embodiment of the present application provides a method for locating the root cause of an alarm, comprising the following steps:

[0049] Step 101: performing cluster analysis on the alarm data set to obtain multiple alarm data subsets;

[0050] In an embodiment of the present application, first, it is necessary to collect alarm data of each fault component from the upstream and downstream systems of the business process. The alarm data includes but is not limited to alarm components, alarm indicators, alarm time, and data related to the alarm system and alarm events, such as business anomaly information, abnormal indicator status, alarm type information, and indicator alarms.

[0051] Then, since the collected alarm data may contain noise, redundant or incomplete information, it is necessary to preprocess the collected alarm data to obtain a preprocessed alarm data set. Optionally, the alarm data set is obtained by preprocessing, and the preprocessing includes at least one of the following:

[0052] Data cleaning; format conversion; feature extraction; data standardization; data normalization.

[0053] It is understandable that through the above preprocessing, duplicate data, invalid data or erroneous data can be removed, data format verification, outlier detection, etc. can be performed.

[0054] Next, a clustering algorithm is used to perform cluster analysis on the alarm data set to obtain multiple alarm data subsets. The embodiment of the present application uses a clustering algorithm to aggregate alarm data with similar alarm patterns into an alarm data subset to facilitate subsequent analysis and processing. Here, the clustering algorithms that can be used include k-means (k-means clustering algorithm) and FCM (fuzzy c-means algorithm, fuzzy c-means clustering algorithm).

[0055] However, since FCM provides flexible clustering, robustness to noise and outliers, and the ability to model fuzzy boundaries, FCM is more suitable for handling situations where ambiguity and overlap exist in alarm data sets than k-means, and can more accurately capture the similarity and correlation between alarm data. Therefore, the preferred embodiment of the present application uses FCM to perform cluster analysis on the alarm data set to obtain multiple alarm data subsets.

[0056] Specifically, the FCM algorithm calculates the fuzzy membership between data points and cluster centers to obtain the membership value of each data point to each cluster center. This fuzzy membership indicates the degree of relationship between the data point and each cluster center. The FCM algorithm is a clustering algorithm based on fuzzy logic. It allows a data point to belong to multiple clusters and assigns a membership to each data point, indicating its degree of membership to each cluster. In the FCM algorithm, each data point is associated with the cluster center through a set of fuzzy membership values. These fuzzy membership values ​​indicate the degree to which the data point belongs to each cluster. The membership value ranges from 0 to 1, indicating the degree of association between the data point and the cluster. The closer to 1, the higher the association.

[0057] In one embodiment, optionally, cluster analysis is performed on the alarm data set to obtain multiple alarm data subsets, including:

[0058] Randomly selecting a plurality of second alarm data from the alarm data set as initial cluster centers;

[0059] For each of the alarm data except the plurality of second alarm data in the alarm data set, obtaining a degree of membership of the alarm data relative to each of the initial cluster centers;

[0060] Iteratively updating the initial cluster center according to the membership degree to obtain multiple cluster centers after iteration;

[0061] A plurality of alarm data sets are obtained based on the plurality of cluster centers.

[0062] In the embodiment of the present application, the goal of the FCM algorithm clustering is to find the optimal membership and cluster center for each alarm data by minimizing the following objective function:

[0063]

[0064] Among them, x i is the i-th data point; c j is the jth cluster center; n is the sample point; c is the cluster center; u ij is the data point x i The membership degree of cluster j; m is the fuzziness factor (usually 2), which controls the fuzziness of the membership degree; ||xi -c j || is the data point x i To cluster center c j The Euclidean distance of .

[0065] Figure 2 Schematic diagram of the FCM algorithm processing process in the embodiment of this application. Figure 2 As shown, the steps of the FCM algorithm are as follows:

[0066] Initialization: Initialize the membership matrix u so that u ij The value of is between 0 and 1 and satisfies

[0067] Calculate cluster centers: Calculate the center of each cluster based on the current membership matrix:

[0068]

[0069] Update the membership matrix: Update the membership matrix according to the new cluster centers:

[0070]

[0071] Check convergence: If the change in the membership matrix is ​​less than a certain threshold, the algorithm terminates; otherwise, return to the step of calculating cluster centers and continue iterating.

[0072] Figure 3 Schematic diagram of the FCM algorithm membership in an embodiment of the present application. Figure 4 A schematic diagram of the center point of the FCM algorithm in an embodiment of the present application; Figure 5 This is a schematic diagram of the FCM algorithm running data in an embodiment of the present application. Figure 6 This is a schematic diagram of the FCM algorithm operation results in the embodiment of this application. Figures 3 to 6 , further details on the FCM algorithm:

[0073] First, it is necessary to generate data that can be used for the FCM algorithm. For the convenience of visualization, a two-dimensional data is generated to facilitate display on the coordinate axis, that is, each sample point and u xj The distance relationship, for example, Figure 3 As shown, u 2j >u 3j , the greater the distance, the smaller the u value. Figure 4 As shown in the figure, the green points in the figure are the cluster centers of these sample points.

[0074] Each sample point includes two features (x coordinate and y coordinate), and 100 such sample points are generated. Of course, the sample points can be changed so that they appear to belong to different classes. Figure 5 shown.

[0075] The steps of the FCM algorithm are as follows:

[0076] (1) Determine the number of classifications, that is, the value of the index m, and the number of iterations (this is the end condition, of course there can be multiple end conditions).

[0077] (2) Initialize a membership degree U (note the condition - the sum is 1);

[0078] (3) Calculate the cluster center C based on U;

[0079] (4) Calculate the objective function J;

[0080] (5) Return to calculate U based on C, return to step (3), and loop until the end, where the loop and end conditions are usually based on the change of the objective function J or reaching a preset number of iterations. The FCM algorithm minimizes the objective function through an iterative optimization process. The objective function is usually defined as the sum of the squares of the weighted distances of all data points to their corresponding cluster centers, where the weights are given by the membership matrix U. The simplest and most common end condition is to reach the preset maximum number of iterations. This means that regardless of whether the objective function is still changing significantly, the algorithm will stop after reaching this number of iterations. This condition ensures that the algorithm does not loop indefinitely; another commonly used end condition is to check whether the change in the objective function J is less than a preset threshold. If the difference between the objective function values ​​of two consecutive iterations is very small (that is, less than a small positive number ∈), it can be considered that the algorithm has converged and the iteration can be stopped. This method relies on the change in the objective function value to determine whether a stable state has been reached. After each iteration, a new objective function value J is calculated. new , and compare it with the objective function value J of the previous iteration old Compare. If |J new -J old If |<∈ (where ∈ is a small positive number), the algorithm is considered to have converged. In this method, we use a combination of the number of iterations and the change in the objective function as the termination condition. That is, the algorithm will first run until it reaches the preset maximum number of iterations. However, if the change in the objective function is sufficiently small (i.e., less than a certain threshold) during the iteration process, the iteration is terminated early. This balances the efficiency and accuracy of the algorithm.

[0081] Step 102: performing frequency analysis on each of the alarm data subsets to obtain a frequency analysis result for each of the alarm data subsets;

[0082] It should be noted that frequency analysis includes frequency analysis and / or number analysis. The embodiment of the present application analyzes the frequency and / or number of occurrences of each alarm data in each of the alarm data subsets to obtain the frequency analysis results of each alarm data set subset. The frequency analysis results include the frequency analysis results and / or number analysis results corresponding to each alarm data.

[0083] In one embodiment, optionally, performing frequency analysis on each of the alarm data subsets to obtain a frequency analysis result corresponding to each of the alarm data subsets includes:

[0084] Using a Gaussian function as a kernel function to obtain a probability density function corresponding to each of the alarm data subsets;

[0085] According to the probability density function, obtaining a first alarm data subset having a probability density greater than a preset probability density in each of the alarm data subsets;

[0086] Perform frequency analysis on the first alarm data subset to obtain a frequency analysis result of each of the alarm data subsets.

[0087] In an embodiment of the present application, for each subset of alarm data, a KDE (kernel density estimation) algorithm is first used for analysis to obtain the probability density function corresponding to each subset of alarm data. Among them, the KDE algorithm is a non-parametric probability density estimation method used to estimate the probability density function of the observations in the data set. It is based on the distribution of sample points, by placing a kernel function around each sample point and superimposing these kernel functions to estimate the probability density distribution in the entire data space. The kernel density algorithm is used to estimate the probability density function of the data. The basic idea is to place a kernel function at each data point, and then superimpose all the kernel functions to obtain a smooth curve, which is the estimated result of the probability density function.

[0088] The kernel density algorithm can also flexibly estimate the shape of the data distribution without pre-assuming the distribution form, making it applicable to a variety of data types and distribution types. It can capture the local characteristics of the data and overcome the limitations of fixed interval widths in methods such as histograms. It estimates the probability density function of the observations in a dataset through the weighted summation of kernel functions. Adjusting the choice of kernel function and bandwidth parameters can flexibly estimate the shape of the data distribution and provide a description of the local characteristics of the data. The KDE algorithm is based on the kernel function and uses a certain bandwidth parameter to estimate the probability density of each alarm data by taking a weighted average of the kernel functions near the alarm data, thereby inferring the population based on a limited data sample.

[0089] Here, the Gaussian function is used as the kernel function because the Gaussian function can map samples into an infinite-dimensional feature space and perform linear classification in this space, thereby effectively handling nonlinear problems. Moreover, the Gaussian function has a smoothing characteristic, with smaller values ​​between samples that are farther apart and larger values ​​between samples that are closer. This smoothing characteristic helps to resist the influence of noise and outliers, and improves the accuracy and robustness of alarm root cause location. The Gaussian function can estimate the overall probability density function through kernel smoothing based on the local density information of the data samples. This ability enables the Gaussian kernel function to more accurately reflect the distribution characteristics of the data in the KDE algorithm, thereby providing a more reliable basis for alarm root cause location. In addition, the Gaussian function can handle a wider range of data distributions, including those data sets with complex nonlinear relationships. The formula of the Gaussian function is as follows:

[0090]

[0091] Here, x is the input value and K(x) represents the value of the Gaussian function.

[0092]

[0093] Where f^(x) is the estimated probability density function at x, n is the sample size, K(x) is the kernel function, and h is the bandwidth.

[0094] It should be noted that the basic idea of ​​the KDE algorithm is to assume that the data points are independently sampled from an unknown probability density function. The goal of the embodiments of this application is to estimate this unknown probability density function. The KDE algorithm does not rely on any specific probability distribution assumptions and is therefore applicable to any type of data.

[0095] Furthermore, based on the probability density function corresponding to each alarm data subset, a first alarm data subset with each probability density greater than a preset probability density is identified. The first alarm data subset can be called a frequent alarm data subset, and frequency analysis is performed based on the first alarm data subset to obtain a frequency analysis result of the alarm data subset to which the first alarm data subset belongs.

[0096] Step 103: establishing a fault propagation graph corresponding to the alarm data set based on the plurality of frequency analysis results, wherein the fault propagation graph includes a plurality of nodes and edges connecting two nodes, wherein the nodes represent the alarm data and the edges represent the causal relationship between the two alarm data;

[0097] In an embodiment of the present application, a fault propagation diagram is a graphical tool for representing the causal relationship between alarm data in upstream and downstream systems of a business process. Figure 7 This is a schematic diagram of the fault propagation diagram described in the embodiment of this application. Figure 7 As shown in the fault propagation graph, nodes represent alarm data, and edges between two nodes represent the causal relationship between the two alarm data corresponding to those nodes—that is, one alarm data may trigger another alarm data. Furthermore, each edge is associated with a confidence level, which indicates the reliability of the causal relationship between the two alarm data. A higher confidence level indicates a stronger causal relationship between the two alarm data. Colors in the graph indicate the priority of the alarm data, and arrows indicate the propagation direction.

[0098] In one embodiment, optionally, establishing a fault propagation graph corresponding to the alarm data set based on the plurality of frequency analysis results includes:

[0099] According to the plurality of frequency analysis results, obtaining a plurality of first alarm data in the alarm data set, the plurality of first alarm data having a number greater than a preset number or a frequency greater than a preset frequency;

[0100] Perform correlation analysis on the plurality of first alarm data sets to establish a fault propagation graph corresponding to the alarm data set.

[0101] In the embodiment of the present application, the first alarm data is alarm data in the alarm data set that appears more than a preset number of times or has an appearance frequency greater than a preset frequency, so the first alarm data can be called frequent alarm data.

[0102] For multiple first alarm data, since at least two of them are often generated together, there is a primary and secondary relationship between these first alarm data. The association analysis includes: mining the primary and secondary relationship of the multiple first alarm data, dividing the multiple first alarm data into primary alarm data and secondary alarm data, and forming triple information, such as Figure 7 As shown, the triplet information includes the primary alarm data, the secondary alarm data corresponding to the primary alarm data, and the confidence between the primary alarm data and the secondary alarm data. The higher the confidence, the stronger the causal relationship between the primary alarm data and the secondary alarm data.

[0103] Furthermore, a fault propagation graph is established based on the mined triple information, where the nodes are alarm data and the edges are confidence levels, which represent the reliability of the causal relationship between two alarms.

[0104] It should be noted that primary and secondary alarm data can be derived from system monitoring and log analysis. Among multiple primary alarm data, one or a few are typically the root cause or primary trigger for other primary alarm data. These primary alarm data are referred to as primary alarm data. Alarm data directly or indirectly triggered by primary alarm data is referred to as secondary alarm data. Secondary alarm data may be a system response to primary alarm data or generated as a result of a chain reaction caused by primary alarm data.

[0105] Step 104 : locating the root cause of the fault propagation graph using a spanning tree principle to determine the root cause corresponding to the alarm data set.

[0106] In the embodiment of the present application, the spanning tree principle is used to locate the root cause of the fault propagation graph and determine the root node of the fault propagation graph. It should be noted that the root node of the fault propagation graph is the root cause corresponding to the alarm data set. In the embodiment of the present application, the root cause can be simply referred to as the root cause.

[0107] In one embodiment, optionally, a spanning tree principle is used to locate a root cause of the fault propagation graph to determine a root cause corresponding to the alarm dataset, including:

[0108] If there is a closed loop in the fault propagation graph, the pruning rule in the spanning tree principle is used to prune the loop and locate the root cause of the fault propagation graph to determine the root cause corresponding to the alarm data set.

[0109] In the embodiment of the present application, the pruning rule in the spanning tree principle is used to prune the fault propagation graph, and the root node is determined according to the fault propagation graph after pruning to locate the root cause. The root node is the root cause.

[0110] It should be noted that closed loops (i.e., the causal relationship between alarm data forms a closed loop) sometimes appear in the fault propagation graph, which generally indicates a misunderstanding or repeated calculation of the fault propagation path. Therefore, the embodiments of the present application use pruning rules based on the spanning tree principle to eliminate these closed loops, thereby simplifying the fault analysis process and improving diagnostic efficiency.

[0111] In the process of pruning the fault propagation graph using the pruning rule in the spanning tree principle, the embodiment of the present application needs to combine the confidence of each edge in the fault propagation graph and give priority to pruning edges with lower confidence.

[0112] Specifically, the pruning rule will find and delete all closed loops in the fault propagation graph, while retaining all nodes and as many edges as possible in the fault propagation graph. Figure 7 As shown, the edges with confidence levels of 3, 6, and 8 are pruned. This is because these three edges form a closed loop. To make the fault propagation graph connected, one of the edges with confidence levels of 7 and 8 needs to be pruned. Both edges have high confidence levels. However, to reduce the complexity of the fault propagation graph and avoid complicating or misleading fault analysis due to the existence of this edge, the edge with confidence level of 8 can also be pruned. The fault propagation graph obtained after pruning the loop is an acyclic graph that can clearly show the main propagation paths between alarm data, thereby helping operation and maintenance personnel quickly locate the root cause of the fault.

[0113] Figure 8This is a flow chart of one implementation of the alarm root cause location method described in the embodiment of this application. Figure 8 As shown, the method includes the following steps:

[0114] Step 801, collecting alarm data sets;

[0115] Step 802: pre-process the alarm data set to obtain a pre-processed alarm data set;

[0116] Step 803: performing cluster analysis on the pre-processed alarm data set to obtain multiple alarm data subsets;

[0117] Step 804: Perform frequency analysis on the subset of alarm data to obtain frequency analysis results;

[0118] Step 805: Perform correlation analysis on the alarm data set based on the frequency analysis results to obtain a fault propagation diagram;

[0119] Step 806: locating the root cause based on the fault propagation diagram;

[0120] Figure 9 This is a flow chart of another embodiment of the alarm root cause location method described in the embodiment of this application. Figure 9 As shown, the method includes the steps shown in the figure, and the description of these steps is omitted here.

[0121] In summary, the alarm root cause location method described in the embodiment of the present application is adopted, the alarm data set is clustered analyzed based on the FCM algorithm, and the frequency analysis is performed in combination with the KDE algorithm. The root cause is located by establishing a fault propagation graph, and the pruning rules in the spanning tree principle are used in the fault propagation graph to improve the efficiency and accuracy of alarm root cause location, thereby reducing manual workload, and being able to adaptively discover new root causes from the constantly changing alarm data, and can be flexibly applied to different types of fault scenarios.

[0122] The present application also provides a device for locating the root cause of an alarm. Figure 10 As shown, including:

[0123] A first obtaining module 1001 is configured to perform cluster analysis on the alarm data set to obtain multiple alarm data subsets;

[0124] The second obtaining module 1002 is configured to perform frequency analysis on each of the alarm data subsets to obtain a frequency analysis result of each of the alarm data subsets;

[0125] Establishing module 1003, configured to establish a fault propagation graph corresponding to the alarm data set based on the plurality of frequency analysis results, wherein the fault propagation graph includes a plurality of nodes and edges connecting two nodes, wherein the nodes represent the alarm data and the edges represent the causal relationship between the two alarm data;

[0126] The determination module 1004 is configured to locate the root cause of the fault propagation graph using a spanning tree principle to determine the root cause corresponding to the alarm data set.

[0127] Optionally, in the alarm root cause locating device, the second obtaining module 1002 is specifically configured to:

[0128] Using a Gaussian function as a kernel function to obtain a probability density function corresponding to each of the alarm data subsets;

[0129] According to the probability density function, obtaining a first alarm data subset having a probability density greater than a preset probability density in each of the alarm data subsets;

[0130] Perform frequency analysis on the first alarm data subset to obtain a frequency analysis result of each of the alarm data subsets.

[0131] Optionally, in the alarm root cause locating device, the establishing module 1003 is specifically configured to:

[0132] According to the plurality of frequency analysis results, obtaining a plurality of first alarm data in the alarm data set, the plurality of first alarm data having a number greater than a preset number or a frequency greater than a preset frequency;

[0133] Perform correlation analysis on the plurality of first alarm data sets to establish a fault propagation graph corresponding to the alarm data set.

[0134] Optionally, in the alarm root cause locating device, the determining module 1004 is specifically configured to:

[0135] If there is a closed loop in the fault propagation graph, the pruning rule in the spanning tree principle is used to prune the loop and locate the root cause of the fault propagation graph to determine the root cause corresponding to the alarm data set.

[0136] Optionally, the alarm root cause locating device performs cluster analysis on the alarm data set to obtain multiple alarm data subsets, including:

[0137] Randomly selecting a plurality of second alarm data from the alarm data set as initial cluster centers;

[0138] For each of the alarm data except the plurality of second alarm data in the alarm data set, obtaining a degree of membership of the alarm data relative to each of the initial cluster centers;

[0139] Iteratively updating the initial cluster center according to the membership degree to obtain multiple cluster centers after iteration;

[0140] A plurality of alarm data sets are obtained based on the plurality of cluster centers.

[0141] Optionally, in the alarm root cause locating device, the alarm data set is obtained by preprocessing, and the preprocessing includes at least one of the following:

[0142] Data cleaning; format conversion; feature extraction; data standardization; data normalization.

[0143] It should be noted that the above-mentioned device provided in the embodiment of the present application can implement all the method steps implemented in the above-mentioned method embodiment and can achieve the same technical effect. The parts and beneficial effects of this embodiment that are the same as those in the method embodiment will not be described in detail here.

[0144] The present application also provides a device for locating the root cause of an alarm. Figure 11 As shown, including:

[0145] Processor 1101, memory 1102, transceiver 1103 and programs or instructions stored on the memory 1102 and executable on the processor 1101; when the processor 1101 executes the programs or instructions, each process of the above-mentioned alarm root cause location method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, they will not be repeated here.

[0146] The transceiver 1103 is configured to receive and send data under the control of the processor 1101 .

[0147] Among them, Figure 11 In the embodiment, the bus architecture may include any number of interconnected buses and bridges, specifically various circuits of one or more processors represented by processor 1101 and memory represented by memory 1102, which are linked together. The bus architecture may also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and are therefore not further described herein. The bus interface provides an interface. The transceiver 1103 may be a plurality of components, i.e., a transmitter and a receiver, providing a unit for communicating with various other devices on a transmission medium. For different user devices, the user interface 1104 may also be an interface capable of connecting external or internal devices as required, and the connected devices include but are not limited to a keypad, a display, a speaker, a microphone, a joystick, etc.

[0148] The processor 1101 is responsible for managing the bus architecture and general processing, and the memory 1102 can store data used by the processor 1101 when performing operations.

[0149] The present application also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the various processes of the above-described alarm root cause location method embodiment and achieves the same technical effects. To avoid repetition, the details are not described here. The computer-readable storage medium may be, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0150] An embodiment of the present application also provides a computer program product, including computer instructions. When the computer instructions are executed by a processor, the various processes of the above-mentioned alarm root cause location method embodiment are implemented, and the same technical effects can be achieved. To avoid repetition, they are not described here.

[0151] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0152] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, and the computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), including a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or network equipment, etc.) to execute the methods described in each embodiment of the present application.

[0153] The embodiments of the present application are described above in conjunction with the accompanying drawings, but the present application is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of this application, ordinary technicians in this field can also make many forms without departing from the purpose of this application and the scope of protection of the claims, all of which are within the protection of this application.

Claims

1. A method for locating the root cause of an alarm, characterized in that: include: Perform cluster analysis on the alarm data set to obtain multiple alarm data subsets; Performing frequency analysis on each of the alarm data subsets to obtain a frequency analysis result for each of the alarm data subsets; Establishing a fault propagation graph corresponding to the alarm data set based on the plurality of frequency analysis results, the fault propagation graph comprising a plurality of nodes and edges connecting two of the nodes, wherein the nodes represent the alarm data and the edges represent a causal relationship between the two alarm data; The root cause of the fault propagation graph is located by using a spanning tree principle to determine the root cause corresponding to the alarm data set.

2. The alarm root cause location method according to claim 1, characterized in that: Performing frequency analysis on each of the alarm data subsets to obtain a frequency analysis result corresponding to each of the alarm data subsets includes: Using a Gaussian function as a kernel function to obtain a probability density function corresponding to each of the alarm data subsets; According to the probability density function, obtaining a first alarm data subset having a probability density greater than a preset probability density in each of the alarm data subsets; Perform frequency analysis on the first alarm data subset to obtain a frequency analysis result of each of the alarm data subsets.

3. The alarm root cause location method according to claim 1, characterized in that: Establishing a fault propagation diagram corresponding to the alarm data set based on the plurality of frequency analysis results includes: According to the plurality of frequency analysis results, obtaining a plurality of first alarm data in the alarm data set, the plurality of first alarm data having a number greater than a preset number or a frequency greater than a preset frequency; Perform correlation analysis on the plurality of first alarm data sets to establish a fault propagation graph corresponding to the alarm data set.

4. The alarm root cause location method according to claim 1, characterized in that: The root cause of the fault propagation graph is located using a spanning tree principle to determine the root cause corresponding to the alarm data set, including: If there is a closed loop in the fault propagation graph, the pruning rule in the spanning tree principle is used to prune the loop and locate the root cause of the fault propagation graph to determine the root cause corresponding to the alarm data set.

5. The alarm root cause location method according to claim 1, characterized in that: Perform cluster analysis on the alarm data set to obtain multiple alarm data subsets, including: Randomly selecting a plurality of second alarm data from the alarm data set as initial cluster centers; For each of the alarm data except the plurality of second alarm data in the alarm data set, obtaining a degree of membership of the alarm data relative to each of the initial cluster centers; Iteratively updating the initial cluster center according to the membership degree to obtain multiple cluster centers after iteration; A plurality of alarm data sets are obtained based on the plurality of cluster centers.

6. The alarm root cause location method according to claim 1, characterized in that: The alarm data set is obtained through preprocessing, wherein the preprocessing includes at least one of the following: Data cleaning; Format conversion; feature extraction; data standardization; data normalization.

7. An alarm root cause location device, characterized in that: include: A first acquisition module is used to perform cluster analysis on the alarm data set to obtain multiple alarm data subsets; a second obtaining module, configured to perform frequency analysis on each of the alarm data subsets to obtain a frequency analysis result of each of the alarm data subsets; An establishment module is used to establish a fault propagation graph corresponding to the alarm data set based on the plurality of frequency analysis results, wherein the fault propagation graph includes a plurality of nodes and an edge connecting two of the nodes, wherein the nodes represent the alarm data and the edges represent a causal relationship between the two alarm data; The determination module is used to locate the root cause of the fault propagation graph using a spanning tree principle to determine the root cause corresponding to the alarm data set.

8. An alarm root cause location device, characterized in that: include: A processor, a memory, and a program or instruction stored in the memory and executable on the processor, wherein the processor implements the alarm root cause locating method according to any one of claims 1 to 6 when executing the program or instruction.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the alarm root cause locating method according to any one of claims 1 to 6 is implemented.

10. A computer program product, characterized in that The method comprises computer instructions, which, when executed by a processor, implements the alarm root cause locating method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Data center anomaly detection method and device based on Gaussian distribution

    CN111737099A

  • Alarm level identification method and device, electronic equipment and storage medium

    CN112100037A

  • Work order filing processing method and device and electronic equipment

    CN114676855A

  • Fault automatic detection and repair method for self-healing intelligent power line

    CN118739184A

  • Fault warning method and device, computer equipment, readable storage medium and program product

    CN119597594A