Alarm root cause positioning method, device and equipment in operation and maintenance scene

By correcting the reputation score in the alarm knowledge graph, using the sigmoid function and dynamic sensitivity adjustment factor, the accuracy of the knowledge graph in the alarm root cause positioning is solved, and the efficiency and accuracy of operation and maintenance work are improved.

CN120455245APending Publication Date: 2025-08-08NEUSOFT CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510757451.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

When using the knowledge graph to locate the alarm, the existing operation and maintenance platform is insufficient in accuracy, which affects the effective execution of operation and maintenance work.

Method used

By correcting the alarm knowledge graph, the sigmoid function is used to assign reputation scores to the alarm standard relationship, and the dynamic sensitivity adjustment factor is determined based on the error frequency data, the reputation score function is adjusted, and the reputation score in the alarm knowledge graph is corrected.

Benefits of technology

It improves the accuracy of the positioning of the root cause of alarms, improves the effectiveness of the knowledge graph in operation and maintenance work, and ensures the accuracy and reliability of the alarm knowledge graph.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120455245A_ABST
    Figure CN120455245A_ABST
Patent Text Reader

Abstract

The invention discloses an alarm root cause positioning method, device and equipment in an operation and maintenance scene. And obtaining an alarm knowledge graph constructed based on the operation and maintenance scene and error statistical data of the alarm knowledge graph. Obtaining a reputation score of an edge of each positioning error root cause in the alarm knowledge graph; a dynamic sensitivity adjustment factor for the corresponding edge is determined based on the error frequency data. And adjusting the reputation scoring function of the corresponding edge by using the dynamic sensitivity adjustment factor to obtain the adjusted reputation scoring function corresponding to the edge of each positioning error root cause. And correcting the reputation score of the edge by using the corresponding adjusted reputation scoring function to obtain the corrected reputation score of the edge of each positioning error root cause so as to correct the alarm knowledge graph. And positioning the alarm root cause of the new alarm data by using the corrected alarm knowledge graph. The alarm knowledge graph corrected by the method can reflect the accuracy performance of the positioning root cause of each side more truly and accurately, so that the accuracy of the positioning of the alarm root cause in an operation and maintenance scene is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of operation and maintenance technology, and in particular to a method, device and equipment for locating the root cause of an alarm in an operation and maintenance scenario. Background Art

[0002] As information systems and network environments become increasingly complex, the types of data that need to be monitored are also becoming increasingly complex. Operations and maintenance engineers often need to monitor large amounts of data to accurately locate the root cause of alarms after receiving them, thereby enabling operations and maintenance management. Currently, some operations and maintenance platforms leverage knowledge graphs to locate the root cause of alarms. However, if the knowledge graphs are inaccurate, this can affect the accuracy of root cause location, hindering the effective execution of operations and maintenance. Summary of the Invention

[0003] Based on the above problems, this application provides a method, device and equipment for locating the root cause of alarms in operation and maintenance scenarios. The purpose is to improve the accuracy of locating the root cause of alarms in operation and maintenance scenarios by correcting the knowledge graph, and to improve the effectiveness of the knowledge graph for operation and maintenance work.

[0004] The embodiments of this application disclose the following technical solutions:

[0005] In a first aspect, the present application provides a method for locating the root cause of an alarm in an operation and maintenance scenario, the method comprising:

[0006] Obtain an alarm knowledge graph; the alarm knowledge graph is constructed based on the operation and maintenance scenario, the entities in the alarm knowledge graph represent alarm types, and the edges between entities represent the link relationship between alarm types;

[0007] Obtain error statistics of the alarm knowledge graph; the error statistics include edges of each root cause of positioning errors within a statistical period and corresponding error frequency data;

[0008] Obtaining a reputation score for each edge in the alarm knowledge graph that locates the root cause of the error; wherein the reputation score is positively correlated with the accuracy of the corresponding edge in locating the root cause;

[0009] determining a dynamic sensitivity adjustment factor of a corresponding edge based on the error frequency data;

[0010] Using the dynamic sensitivity adjustment factor, adjust the reputation scoring function of the corresponding edge to obtain an adjusted reputation scoring function corresponding to each edge of the root cause of the positioning error;

[0011] For each edge of the root cause of the positioning error, correct the reputation score using the corresponding adjusted reputation scoring function to obtain a corrected reputation score for the edge of each root cause of the positioning error, so as to correct the alarm knowledge graph;

[0012] For new alarm data, the root cause of the alarm is located using the corrected alarm knowledge graph.

[0013] A second aspect of the present application provides a device for locating the root cause of an alarm in an operation and maintenance scenario, the device comprising:

[0014] A graph acquisition module is used to obtain an alarm knowledge graph; the alarm knowledge graph is constructed based on the operation and maintenance scenario, the entities in the alarm knowledge graph represent alarm types, and the edges between entities represent the link relationship between alarm types;

[0015] A data acquisition module is used to obtain error statistics of the alarm knowledge graph; the error statistics include edges of each root cause of positioning errors within a statistical period and corresponding error frequency data;

[0016] A score acquisition module is used to obtain the reputation score of each edge in the alarm knowledge graph that locates the root cause of the error; the reputation score is positively correlated with the accuracy of the corresponding edge in locating the root cause;

[0017] an adjustment factor determination module, configured to determine a dynamic sensitivity adjustment factor of a corresponding edge based on the error frequency data;

[0018] A function adjustment module, configured to adjust the reputation scoring function of the corresponding edge using the dynamic sensitivity adjustment factor to obtain an adjusted reputation scoring function corresponding to each edge of the root cause of the positioning error;

[0019] A graph correction module, configured to correct the reputation score of each edge of each root cause of the positioning error using the corresponding adjusted reputation scoring function to obtain a corrected reputation score of each edge of the root cause of the positioning error, so as to correct the alarm knowledge graph;

[0020] The root cause location module is used to locate the root cause of alarms based on new alarm data using the corrected alarm knowledge graph.

[0021] A third aspect of the present application provides a device for locating the root cause of an alarm in an operation and maintenance scenario, the device comprising: a processor and a memory communicatively connected to each other;

[0022] The memory stores a computer program;

[0023] The processor is used to run the computer program to implement the alarm root cause location method in the operation and maintenance scenario introduced in any implementation manner of the first aspect.

[0024] Compared with the existing technology, this application has the following beneficial effects:

[0025] The method for locating the root cause of an alarm in an operation and maintenance scenario provided by the present application requires obtaining an alarm knowledge graph constructed based on the operation and maintenance scenario, in which entities represent alarm types, and edges between entities represent link relationships between alarm types. In addition, error statistics of the alarm knowledge graph are obtained, which include edges of each root cause of the located error during the statistical period and corresponding error frequency data. The reputation score of each edge of the root cause of the located error in the alarm knowledge graph is obtained; and the dynamic sensitivity adjustment factor of the corresponding edge is determined based on the error frequency data. Then, the reputation scoring function of the corresponding edge is adjusted using the dynamic sensitivity adjustment factor to obtain the adjusted reputation scoring function corresponding to the edge of each root cause of the located error. For each edge of the root cause of the located error, the reputation score is corrected using the corresponding adjusted reputation scoring function to obtain the corrected reputation score of the edge of each root cause of the located error, so as to correct the alarm knowledge graph. For new alarm data, the root cause of the alarm is located using the corrected alarm knowledge graph. In the technical solution of the present application, based on the error frequency data of the edge for locating the root cause of the error, a personalized adjustment of the reputation scoring function of the edge is achieved through a dynamic sensitivity adjustment factor, and then the reputation score of the edge is adjusted using the adjusted reputation scoring function to achieve the correction of the alarm knowledge graph. Since the reputation scores of some edges in the corrected alarm knowledge graph have been adjusted, the accuracy of the root cause of each edge is reflected more realistically and accurately, so that the corrected alarm knowledge graph can more accurately assist in locating the root cause of the alarm for new alarm data. It can be seen that this method can improve the accuracy of locating the root cause of the alarm in the operation and maintenance scenario by correcting the knowledge graph, and improve the effectiveness of the knowledge graph for operation and maintenance work. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0027] Figure 1A This is a schematic diagram of the morphology of the sigmoid function curve;

[0028] Figure 1B A flowchart of a method for locating the root cause of an alarm in an operation and maintenance scenario provided in an embodiment of the present application;

[0029] Figure 1C A diagram of the training architecture of an alarm classification model provided in an embodiment of the present application;

[0030] Figure 1D This is a structural diagram of an alarm knowledge graph;

[0031] Figure 2A This is a diagram of the implementation architecture of a correction alarm knowledge graph in an operation and maintenance scenario provided by an embodiment of the present application;

[0032] Figure 2B A flowchart of a method for locating the root cause of an alarm in another operation and maintenance scenario provided by an embodiment of the present application;

[0033] Figure 2C A flowchart of another method for locating the root cause of an alarm in an operation and maintenance scenario provided in an embodiment of the present application;

[0034] Figure 3 A schematic diagram of the construction process of an alarm knowledge graph based on ensemble learning provided in an embodiment of the present application;

[0035] Figure 4 A schematic diagram of a process for locating the root cause of an alarm provided in an embodiment of the present application;

[0036] Figure 5 A schematic diagram of an embodiment of the present application providing an offline-online hybrid correction mechanism for updating a knowledge graph to locate the root cause of an alarm;

[0037] Figure 6 This is an implementation architecture diagram for locating the root cause of an alarm using a multi-dimensional knowledge graph, provided in an embodiment of the present application;

[0038] Figure 7 A schematic diagram of the structure of an alarm root cause locating device in an operation and maintenance scenario provided in an embodiment of the present application. DETAILED DESCRIPTION

[0039] As previously described, when using knowledge graphs to locate the root cause of network monitoring alarms, the lack of accuracy in the knowledge graphs often leads to inaccurate root cause location, which in turn interferes with operational maintenance work related to the alarms.

[0040] To address the above issues, the inventors propose in this application that the sigmoid function can be used to adjust the knowledge graph in operation and maintenance scenarios. The sigmoid function is a commonly used mathematical function that has the form and characteristics of an S-shaped curve, and the function's value range is between 0 and 1. The function expression is as follows:

[0041] ; Formula (1)

[0042] On the function curve of the sigmoid function, the middle part of the curve changes quickly, while the two sides change slowly, such as Figure 1AAs shown in Figure 2, it shows the shape of the sigmoid function curve. This characteristic of the sigmoid function can be used to perform reputation scoring on edges between entities in the knowledge graph, using the reputation score to reflect the credibility and importance of the edge in the knowledge graph.

[0043] In operations and maintenance scenarios, alarm types can be considered entities in the alarm knowledge graph, and the relationship between two alarm types can be considered an edge in the alarm knowledge graph. The relationship between alarm types can also be referred to as the alarm standard relationship in the alarm knowledge graph. Using the sigmoid function to assign a reputation score to the alarm standard relationship reflects its credibility and importance.

[0044] When locating the root cause of an alarm, if a standard relationship for an alarm is found to be incorrect, its reputation score is lowered; if a standard relationship for an alarm is found to be correct, its residual score is increased. This effectively penalizes incorrect information while gradually increasing the residual score of correct information, ensuring the overall quality of the alarm knowledge graph.

[0045] However, as shown in formula (1) above, the coefficient in front of the independent variable of the sigmoid function is fixed, which means that the steepness of the function (or the slope of the function) is fixed and cannot be adjusted. It should be noted that the fixed and unadjustable steepness here does not mean that the slope at each point on this function curve is completely uniform, but rather that the scoring method remains constant when the sigmoid function is used to score the reputation of each edge in the alarm knowledge graph. This leads to a problem: past error statistics of the alarm knowledge graph, such as the error frequency data of the edge that locates the root cause of the error, do not have a different impact on the penalty (i.e., score reduction) of the edge's reputation score. In other words, whether the error frequency is high or low, there is no difference in the reputation score of the error edge. Therefore, the accuracy of the alarm knowledge graph is affected.

[0046] To correct the accuracy of the alarm knowledge graph and thereby improve the accuracy of alarm root cause location, the inventors propose in this application to further improve the sigmoid function. Specifically, they determine a dynamic sensitivity adjustment factor for edges in the alarm knowledge graph based on error statistics from the alarm knowledge graph. This factor is used to adjust the reputation scoring function (i.e., the sigmoid function). Consequently, the reputation score for each edge is recalibrated using the adjusted reputation scoring function. This achieves the goal of correcting the alarm knowledge graph and improving the accuracy of alarm root cause location.

[0047] In order to help those skilled in the art better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.

[0048] See also Figure 1B , which is a flow chart of a method for locating the root cause of an alarm in an operation and maintenance scenario provided by an embodiment of the present application. Figure 1B The method for locating the root cause of an alarm in the operation and maintenance scenario shown includes the following steps:

[0049] S101. Obtain an alarm knowledge graph.

[0050] The alarm knowledge graph is constructed based on operational scenarios. Entities in the alarm knowledge graph represent alarm types, and edges between entities represent the links between alarm types. Alarm types are determined by classifying historical alarm data. Alarm associations are discovered and their root causes are located, forming links between alarm types and ultimately forming the alarm knowledge graph. This step obtains the constructed alarm knowledge graph so that it can be corrected in subsequent steps.

[0051] Generally speaking, when abnormal indicators exceed limits, failures occur, or system performance degrades during system operation, these can impact system stability and business continuity. In these cases, the system can output alarm data to communicate these critical issues to monitoring equipment or personnel responsible for monitoring system alarms. In practical applications, the alarm type represented by each entity node in the alarm knowledge graph can be determined using an alarm classification model.

[0052] In an example implementation, the present application can use a supervised learning approach to pre-build an alarm classification model, and then use the trained alarm classification model to complete the classification of alarm data. Before building the model, the alarm data can be manually classified according to the definition of key indicator alarms and alarm categories of each business system resource to form a three-dimensional alarm classification system such as a "resource-indicator-abnormality description" structure. The historical alarm data is categorized and used as data input for the alarm classification model. In the above-mentioned three-dimensional alarm classification system, "resources" refer to the system components of the alarm source, such as servers, databases, network equipment, etc.; "indicators" refer to the specific performance parameters monitored, such as CPU utilization, memory usage, disk IO, etc.; "abnormality description" refers to the abnormal performance of the alarm compared to the normal state, such as CPU usage exceeding 90%, or network packet loss rate exceeding the threshold, etc.

[0053] Figure 1C This is a training architecture diagram of an alarm classification model provided in an embodiment of the present application. Figure 1C As shown, the distributed computing service can be used to complete the model training. Specifically, the distributed computing service can perform data preprocessing, such as data cleaning, on the alarm data based on the acquired alarm data and alarm category labels. Afterwards, the alarm content in the alarm data can be segmented, word vectors are calculated, and other operations can be performed before model training. In an optional implementation method, a variety of neural network models with different structures can be used for model training and optimization. By evaluating the performance of the model in terms of prediction accuracy, recall rate, computational efficiency, etc., the best model is selected as the alarm classification model, and the model is stored. Thus, the alarm classification model is used to classify the alarm data when necessary.

[0054] S102. Obtain error statistics of the alarm knowledge graph.

[0055] In the technical solution provided in the embodiments of this application, the alarm knowledge graph can be used to locate the root cause of alarm data. In the embodiments of this application, locating the alarm root cause essentially involves analyzing the causal relationship between the entity nodes related to the alarm in the knowledge graph. However, there are two possible verification results for locating the alarm root cause: one verification result is that the root cause is incorrectly located, and the other verification result is that the root cause is correctly located.

[0056] In this embodiment of the present application, in order to more accurately correct the acquired knowledge graph, it is proposed that error statistics of the alarm knowledge graph be obtained through this step. The error statistics include the edges that locate the root cause of each error within the statistical period and the corresponding error frequency data. For example, the error frequency data of the edges that locate the root cause of the error in the alarm knowledge graph over the past three days is counted. Figure 1D The structure of an alarm knowledge graph is shown as an example. Figure 1DIn the figure, 10 entity nodes are shown as an example, namely entity node A, entity node B...entity node J, which represent 10 different alarm types. The edge between the entity nodes represents the link relationship between the two alarm types represented by the two connected entity nodes. This link relationship can specifically refer to a causal relationship under the alarm root cause location requirement in the operation and maintenance scenario. The following is an example of error frequency data: as an example, in the alarm knowledge graph in the past three days, the frequency of edge AB between entity node A and entity node B locating the root cause of the error is 6 times, and the frequency of edge AC between entity node A and entity node C locating the root cause of the error is 8 times. Similarly, the edges of each root cause of the error location in the alarm knowledge graph and its error frequency data during the statistical period can be obtained. The error frequency data will be used to determine the dynamic sensitivity adjustment factor of the edge in step S104 described later. For details, please refer to the introduction below. In addition, in order to correct the accuracy of the alarm knowledge graph in this application, it is also necessary to obtain the reputation score of the edge of the root cause of the error location through step S103.

[0057] S103. Obtain the reputation score of each edge of the root cause of the positioning error in the alarm knowledge graph.

[0058] In this embodiment of the present application, the reputation score of an edge is positively correlated with the accuracy of the corresponding edge-located root cause. Specifically, the higher the accuracy of the edge-located alarm root cause, the higher its reputation score; the lower the accuracy of the edge-located alarm root cause, the lower its reputation score.

[0059] The reputation scoring system incorporates both penalty and reward mechanisms. The sigmoid function-based reputation scoring system places particular emphasis on penalizing incorrect root causes. Because the sigmoid function grows rapidly in the middle, when a relationship between alarm types (i.e., an edge in the alarm knowledge graph) is verified as locating an incorrect root cause, its reputation score decreases rapidly. This design aligns with the logic of reputation growing slowly but declining rapidly, effectively preventing the accumulation and spread of incorrect information within the knowledge graph. In line with the penalty mechanism, the reputation scoring system also rewards correct root causes. Because the sigmoid function grows slowly on both sides, when a relationship between alarm types (i.e., an edge in the alarm knowledge graph) is verified as locating a correct root cause, its reputation score gradually increases. This design encourages the accumulation and dissemination of correct information, helping to improve the overall quality of the knowledge graph.

[0060] This step obtains the reputation score of the edge in the alert knowledge graph that locates the root cause of the error. This means obtaining the current, latest reputation score of the edge that locates the root cause of the error. This score is determined by the accuracy of the corresponding edge in locating the root cause in the past.

[0061] S104: Determine a dynamic sensitivity adjustment factor of a corresponding edge based on the error frequency data.

[0062] In the embodiment of the present application, as a modification of the sigmoid reputation scoring function, a dynamic sensitivity adjustment factor for the edges in the alert knowledge graph is added. The expression of the dynamic sensitivity adjustment factor k is as follows:

[0063] ; Formula (2)

[0064] In the above formula, γ represents the error sensitivity coefficient; Error_Count represents the error frequency data, which is obtained in step S102; k0 represents the initial sensitivity adjustment factor of the edge, which is a constant value.

[0065] In formula (2), the core factor determining the dynamic sensitivity adjustment factor k is the error frequency data Error_Count. According to formula (2), the larger the value of the error frequency data Error_Count, the larger the value of k; the smaller the value of the error frequency data Error_Count, the smaller the value of k. Because the error frequency data of each edge locating the root cause of the error may be different, the k value for each edge locating the root cause of the error may also be different. Formula (2) makes the dynamic sensitivity adjustment factor of each edge strongly correlated with the error frequency data of that edge.

[0066] S105 : Using the dynamic sensitivity adjustment factor, adjust the reputation scoring function of the corresponding edge to obtain the adjusted reputation scoring function corresponding to each edge of the root cause of the positioning error.

[0067] The dynamic sensitivity adjustment factor k calculated in the previous step S104 is used to adjust the reputation scoring function of the corresponding edge. The adjusted function expression is as follows:

[0068] ; Formula (3)

[0069] By comparing Formula (3) with Formula (1), we can see the specific changes in the reputation scoring function. A dynamic sensitivity adjustment factor k is added before the independent variable x. In Formula (3), k can be used to adjust the steepness of the function, thereby dynamically adjusting the sensitivity of punishment or reward.

[0070] Specifically, if k>1, the reputation scoring function will impose a greater penalty on the edge that locates the root cause of the error, achieving a rapid reduction in the score of the erroneous edge; in addition, the reward will be more gradual, that is, the correct edge will be slowly rewarded. If k<1, the reputation scoring function will impose a more gentle penalty on the edge that locates the root cause of the error; in addition, the reward will be greater, and the correct edge will be rewarded more quickly. It can be understood that if k>1, the sensitivity to the penalty of the erroneous edge is increased; if k<1, the sensitivity to the penalty of the erroneous edge is reduced. In practical applications, combined with formula (2), the γ value in formula (2) can be set according to the sensitivity requirements in specific operation and maintenance scenarios. In addition, combined with formula (2) and formula (3), it can be seen that the error frequency data of the edge that locates the root cause of the error differentially affects the value of the dynamic sensitivity adjustment factor of the corresponding edge, and then the reputation scoring function of the corresponding edge is adjusted in a targeted manner, so that the reputation scoring function can make the correction effect more accurate and personalized when correcting the reputation score of the corresponding edge, thereby ensuring the accuracy of the graph correction effect.

[0071] S106. For each edge of the root cause of the positioning error, correct the reputation score using the corresponding adjusted reputation scoring function to obtain the corrected reputation score of the edge of each root cause of the positioning error, so as to correct the alarm knowledge graph.

[0072] In an optional implementation, for each edge that is the root cause of the positioning error, before using the corresponding adjusted reputation scoring function to correct the reputation score, the independent variable increment of the reputation scoring function is also obtained. For the same edge, the relationship between the old and new reputation scores is expressed as follows:

[0073] ; Formula (4)

[0074] In formula (4), The increment of the independent variable of the edge that represents the root cause of a positioning error, where α is the penalty coefficient and |E| represents the severity of the error. represents the reputation score before the edge correction, represents the newly obtained reputation score after the edge is corrected. σ( ) is the adjusted reputation score function corresponding to the edge, as shown in Formula (3).

[0075] By applying Formula (4) in the above example, we can calculate the new reputation score of a certain edge with a known root cause of a positioning error based on its current reputation score, penalty coefficient, and error severity, using the adjusted reputation scoring function of Formula (3). This new reputation score is used as the corrected reputation score to replace the reputation score of the edge in the alarm knowledge graph before this correction. Similarly, the reputation scores of the edges with each root cause of positioning errors recorded during the statistical period are corrected one by one in this way, completing the correction of the entire alarm knowledge graph.

[0076] S107: For new alarm data, locate the root cause of the alarm using the corrected alarm knowledge graph.

[0077] As described above, by adjusting the reputation scoring function of edges in the alarm knowledge graph that locate error root causes using a dynamic sensitivity adjustment factor, we can calibrate edge reputation scores based on error frequency data, thereby correcting the entire alarm knowledge graph. After completing this calibration of the alarm knowledge graph, when new alarm data arrives, the calibrated alarm knowledge graph can be used to locate the root cause of the new alarm data. This improves the accuracy of locating the root cause of new alarm data.

[0078] The method for locating the root cause of an alarm in an operation and maintenance scenario provided by the present application requires obtaining an alarm knowledge graph constructed based on the operation and maintenance scenario, in which entities represent alarm types, and edges between entities represent link relationships between alarm types. In addition, error statistics of the alarm knowledge graph are obtained, which include edges of each root cause of the located error during the statistical period and corresponding error frequency data. The reputation score of each edge of the root cause of the located error in the alarm knowledge graph is obtained; and the dynamic sensitivity adjustment factor of the corresponding edge is determined based on the error frequency data. Then, the reputation scoring function of the corresponding edge is adjusted using the dynamic sensitivity adjustment factor to obtain the adjusted reputation scoring function corresponding to the edge of each root cause of the located error. For each edge of the root cause of the located error, the reputation score is corrected using the corresponding adjusted reputation scoring function to obtain the corrected reputation score of the edge of each root cause of the located error, so as to correct the alarm knowledge graph. For new alarm data, the root cause of the alarm is located using the corrected alarm knowledge graph. In the technical solution of the present application, based on the error frequency data of the edge for locating the root cause of the error, a personalized adjustment of the reputation scoring function of the edge is achieved through a dynamic sensitivity adjustment factor, and then the reputation score of the edge is adjusted using the adjusted reputation scoring function to achieve the correction of the alarm knowledge graph. Since the reputation scores of some edges in the corrected alarm knowledge graph have been adjusted, the accuracy of the root cause of each edge is reflected more realistically and accurately, so that the corrected alarm knowledge graph can more accurately assist in locating the root cause of the alarm for new alarm data. It can be seen that this method can improve the accuracy of locating the root cause of the alarm in the operation and maintenance scenario by correcting the knowledge graph, and improve the effectiveness of the knowledge graph for operation and maintenance work.

[0079] Figure 2A This is a diagram of the implementation architecture of the correction alarm knowledge graph in an operation and maintenance scenario provided by an embodiment of the present application. Figure 2A The following describes three adjustment strategies for correcting the alarm knowledge graph. Strategies 1 has already been described in the previous examples. The following focuses on the other two adjustment strategies (Strategies 2 and 3).

[0080] In the past, when using reputation scoring functions to score the reputation of edges in the alert knowledge graph, the old feedback and the new feedback have the same influence weight on the reputation of the edge. This scoring method is difficult to reflect the timeliness. For this reason, this application proposes the following Figure 2A Adjustment strategy 2 in this section obtains the time interval between reputation score corrections to determine the time decay factor. This time decay factor is then incorporated into the function's independent variable as a coefficient multiplied by the old reputation score. Furthermore, when applying the adjusted reputation scoring function to correct the edge's reputation score, the time decay factor in the independent variable is also considered, influencing the calculation of the reputation score. This also achieves differentiated treatment of old and new feedback. This is described below in conjunction with the following examples.

[0081] Figure 2B This is a flow chart of another alarm root cause location method in an operation and maintenance scenario provided by an embodiment of the present application. Figure 2B The method shown includes:

[0082] S201 to S205. The implementation of S201 to S205 is substantially the same as that of S101 to S105, and reference may be made to the introduction of the above embodiment, which will not be repeated here.

[0083] S206: Obtain the time interval between the most recent correction time and the current time of the reputation score of each edge of the positioning error root cause.

[0084] It should be noted that the edge of the root cause of the positioning error determined in each statistical period may be different. Figure 1D For example, in the statistical period T1 closest to the current moment, the edges identified as the root cause of a positioning error include edges AB, EF, and HJ. In the statistical period T2 preceding T1, the edges identified as the root cause of a positioning error include edges AB, AE, and EF. In the statistical period T3 preceding T2, the edges identified as the root cause of a positioning error include edges EF and HJ. In this example, since edge AB was identified as the root cause of a positioning error during statistical period T2, its reputation score was corrected once before statistical period T1. Accordingly, since edge EF was not identified as the root cause of a positioning error during statistical period T2 but was identified as the root cause of a positioning error during statistical period T3, its reputation score was corrected once before statistical period T2. In other words, the time interval t1 between the time of the last correction of edge EF's reputation score and the current moment is greater than the time interval t2 between the time of the last correction of edge AB's reputation score and the current moment.

[0085] In conjunction with the introduction to this step, by analyzing the time of the most recent correction of the reputation score of each edge that is the root cause of the positioning error, as included in the error statistics obtained in step S202, the time interval between each correction time and the current time can be calculated. This time interval indicates that the previous correction is old feedback. Obviously, in the above example, since time interval t1 > time interval t2, the correction of the reputation score of edge EF before statistical period T2 is relatively old feedback, while the correction of the reputation score of edge AB before statistical period T1 is relatively new feedback.

[0086] S207: Obtain a time attenuation factor of each edge of the root cause of the positioning error according to each time interval.

[0087] In one example, the time decay factor may be exponential, as shown in the following expression:

[0088] ; Formula (5)

[0089] In formula (5), t represents the time interval. For the calculation method, see step S206. It is a parameter used to control the time decay rate, with a value between 0 and 1.

[0090] It should be noted that the above formula (5) is only an example of a calculation method for the time decay factor. In practical applications, there are no restrictions on the expression for calculating the time decay factor, as long as the time decay factor is negatively correlated with the time interval. In other words, the larger the time interval, the smaller the time decay factor, and the smaller the impact on the reputation score correction; the smaller the time interval, the larger the time decay factor, and the greater the impact on the reputation score correction.

[0091] Next, step S208 is used to introduce a specific implementation of correcting the reputation score of each edge of the root cause of the positioning error using the corresponding adjusted reputation scoring function in this embodiment.

[0092] S208. For each edge of the root cause of the positioning error, obtain the product of the corresponding time attenuation factor and the reputation score, and based on the product, use the corresponding adjusted reputation scoring function to correct the reputation score to obtain the corrected reputation score of the edge of each root cause of the positioning error to correct the alarm knowledge graph.

[0093] ; Formula (6)

[0094] In formula (6), The increment of the independent variable of the edge representing the root cause of a location error. is the time attenuation factor, which can be calculated using formula (5). represents the reputation score before the edge correction, represents the newly obtained reputation score after the edge is corrected. σ( ) is the adjusted reputation score function corresponding to the edge, as shown in Formula (3).

[0095] S209: For new alarm data, locate the root cause of the alarm using the corrected alarm knowledge graph.

[0096] In the above embodiments, the Figure 2A The adjustment ideas 1 and 2 shown not only adjust the structure of the reputation scoring function but also add a time decay factor as a multiplication coefficient to the reputation score before correction. In this embodiment, to achieve accurate correction of the alarm knowledge graph, the influence of the time interval of the correction feedback is taken into account, while also paying attention to the differentiated performance of the error frequency data of the edge that locates the root cause of the error. Compared with the sigmoid reputation scoring function shown in formula (1) and the function independent variable shown in formula (4), the time factor and the error frequency are jointly incorporated into the accuracy of the alarm knowledge graph. Through the implementation of the steps of the embodiment, the accuracy performance of the alarm knowledge graph is effectively corrected. In turn, a more accurate root cause location service is provided for new alarm data.

[0097] The inventors found through research that when calculating the reputation score of the edge in the knowledge graph, the traditional linear increment It is difficult to distinguish the severity of the error. Therefore, when calculating the reputation score using the reputation scoring function for edges with different error levels, the independent variable increment is used. , it is impossible to reflect the difference in the severity of the error through reputation scores. To this end, the inventors propose the following embodiments, see Figure 2C , which is a flow chart of another method for locating the root cause of an alarm in an operation and maintenance scenario provided by an embodiment of the present application. Figure 2C As shown, the method includes:

[0098] S301 to S305. The implementation of S201 to S205 is substantially the same as that of S101 to S105, and reference may be made to the introduction of the above embodiment, which will not be repeated here.

[0099] S306. For each edge that is a root cause of a positioning error, identify the error level of each edge. For edges with a first error level, determine the independent variable increment of the adjusted reputation scoring function by taking the logarithm of the error severity. For edges with a second error level, determine the independent variable increment of the adjusted reputation scoring function by multiplying the error severity by itself. For each edge that is a root cause of a positioning error, use the corresponding adjusted reputation scoring function to correct the reputation score to obtain a corrected reputation score for each edge that is a root cause of a positioning error.

[0100] The first error level is higher than the second error level. The first error level can be considered a major error, while the second error level can be considered a routine error. In practical applications, various possible error types can be classified based on their severity, with some being classified as the first error level and others as the second error level.

[0101] In this step, a piecewise incremental function is defined to distinguish between normal errors and major errors. See formula (7):

[0102] ; Formula (7)

[0103] In formula (7), the first row corresponds to the independent variable increment for calculating the second error level edge; the second row corresponds to the independent variable increment for calculating the first error level edge; and the third row corresponds to the independent variable increment for calculating the edge that locates the correct root cause. This step focuses on the first two rows in formula (7).

[0104] In formula (7), α represents the penalty coefficient and |E| represents the severity of the error. Represents the reward coefficient, and |C| represents the correct contribution. Among them, the penalty coefficient α can be adjusted according to the specific application scenario to control the extent of the reputation score decline; the reward coefficient Adjustments can be made based on specific application scenarios to control the extent to which the reputation score increases. Formula (7) shows that for major errors, a square penalty is triggered, which quickly reduces the reputation score; for correct contributions, a square root function is used to avoid excessive inflation of the reputation score.

[0105] S307. For each edge with the correct root cause located within the statistical period, determine the independent variable increment of the reputation scoring function by taking the square root of the correct contribution; obtain the reputation score of each edge with the correct root cause located in the alarm knowledge graph; for each edge with the correct root cause located within the statistical period, use the initial sensitivity adjustment factor to adjust the reputation scoring function of the corresponding edge to correct the reputation score of the corresponding edge.

[0106] Regarding the implementation of finding the square root of the correct contribution, please refer to the expression shown in the third row of formula (7). The embodiment of the present application not only corrects the reputation score of the edge that locates the wrong root cause, but also corrects the reputation score of the edge that locates the correct root cause. In this step, similar to the previous embodiment, the remaining score of each edge that locates the correct root cause within the statistical period is also obtained. For the reputation scoring function, the initial sensitivity adjustment factor k0 is used to adjust the inherent function, such as the sigmoid function shown in formula (1), with the initial sensitivity adjustment factor k0 of the edge as the coefficient. That is, for the edge that locates the correct root cause, k in formula (3) takes the value k0. In this way, the independent variable increment of each edge that locates the correct root cause and the expression of the adjusted reputation scoring function can be determined, and then the corrected reputation score of the edge that locates the correct root cause can be calculated.

[0107] After the above steps, the reputation scores of the edges that correctly locate the root cause and the edges that incorrectly locate the root cause have been corrected. This completes the correction of the entire alarm knowledge graph. Next, step S308 can be used to more accurately locate the root cause of the alarm using the corrected alarm knowledge graph.

[0108] S308: For new alarm data, locate the root cause of the alarm using the corrected alarm knowledge graph.

[0109] In the above embodiments, the Figure 2A Adjustment Idea 1 + Adjustment Idea 3 shown not only adjusts the structure of the reputation scoring function, but also proposes a hierarchical penalty and reward mechanism for the function's independent variable increment, achieving a nonlinear design for the independent variable increment. In this embodiment, to achieve accurate correction of the alarm knowledge graph, both the severity of the error and the correct contribution are taken into account, while also paying attention to the differentiated performance of the error frequency data of the edge that locates the root cause of the error. Compared with the sigmoid reputation scoring function shown in formula (1) and the function independent variable shown in formula (4), the severity of the error and the error frequency are both incorporated into the accuracy factors of the alarm knowledge graph. Through the implementation of the steps of the embodiment, the accuracy performance of the alarm knowledge graph is effectively corrected. In turn, a more accurate root cause location service is provided for new alarm data.

[0110] It should be noted that the above Figure 2B and Figure 2C The embodiments shown can also be combined with each other. Figure 2AThe three adjustment strategies shown can be combined. Specifically, they consider both the severity of errors and the contribution of correct answers (adjustment strategy 3), the differential performance of error frequency data for edges that identify the root cause of errors (adjustment strategy 1), and the influence of the time interval between correction feedback (adjustment strategy 2). In summary, the sigmoid-based reputation scoring system ensures efficient correction and optimization of the knowledge graph by quickly penalizing root causes of errors and slowly rewarding correct root causes. Among the improvements proposed in this solution, a dynamic sensitivity adjustment factor adapts to the needs of different scenarios, ensuring fine-grained control of the scoring system; a hierarchical mechanism distinguishes error severity, improving the system's noise tolerance; and a time decay factor enhances the real-time impact of score correction. This system enables the knowledge graph to maintain its high quality and reliability, adapting to changing environments and complex application scenarios.

[0111] In the previous embodiments, the alarm knowledge graph was mentioned many times. The following describes the construction method of the alarm knowledge graph in conjunction with the embodiments. In the embodiment of this application, the alarm knowledge graph can be constructed by using an integrated learning method. Specifically:

[0112] First, the historical alarm data is constructed into multiple causal discovery samples according to time. For example, a causal discovery sample is constructed with every 5 minutes as a time window. Of course, the length of the time window can also be determined according to the specific business, which is not limited here. Thereafter, the causal discovery samples are grouped based on the correlation characteristics of the causal discovery samples. Among them, the causal discovery samples whose correlation characteristics meet the preset conditions are divided into the first group, and the causal discovery samples that cannot meet the preset conditions are divided into the second group. The preset conditions include: there is a dependency relationship between the businesses, and / or the dynamic time warping distance after regularization is less than the preset distance threshold. In the following text, the implementation method of grouping causal discovery samples will be explained based on the different contents of the preset conditions.

[0113] Next, the first set of causal discovery samples is analyzed using multiple different base learners to generate a first knowledge graph; and the second set of causal discovery samples is analyzed using multiple different base learners to generate a second knowledge graph. In both the first and second knowledge graphs, entities represent alarm types, and edges between entities represent link relationships between alarm types. That is, the meanings of entities and edges in the first and second knowledge graphs are the same as those in the alarm knowledge graph.

[0114] Finally, the first knowledge graph and the second knowledge graph are unioned to obtain the alarm knowledge graph.

[0115] The first group mentioned above can be called correlation estimation grouping, and the second group can be called random grouping. Among them, correlation estimation grouping is suitable for clear and complete resource links, and the historical alarm data generated by them are attributed to the same sample group. Correlation estimation grouping is suitable for components with clear upstream and downstream dependencies, such as database-application server-front-end interface components. This type of component with clear upstream and downstream dependencies can directly infer causal relationships based on existing system topology, log analysis, historical data, etc. Random grouping is suitable for resources whose correlation relationships cannot be evaluated, and their alarm data are grouped together. It is suitable for alarm data whose correlation cannot be directly inferred. For example, independent nodes in a distributed environment.

[0116] Two optional implementations of grouping causal discovery samples are provided below.

[0117] Implementation method 1: Perform alarm correlation analysis based on business rules to determine sample groups. This method is also called rule matching.

[0118] For example, a causal discovery sample involves two components, A and B, and there is a clear business dependency between components A and B (such as a web server requires database support). It is considered that there is a potential causal relationship between components A and B, and the causal discovery sample can be classified into the first group.

[0119] Implementation method 2: Determine sample grouping based on the Dynamic Time Warping (DTW) method of time series.

[0120] This method uses the DTW algorithm to calculate the similarity of alarm data from different time periods. If the alarm time pattern of a component is consistent with that of its dependent components, they are placed in the same group. The specific steps are as follows: First, the DTW distance between the two alarm sequences is calculated and regularized. Next, if the regularized DTW distance is ≤ a set threshold (for example, 5 minutes), the two alarm sequences have similar time patterns and can be placed in the correlation estimation group, i.e., the first group. If the DTW distance is too large, it indicates that the time patterns of the two sequences differ significantly, and they are added to a random group, i.e., the second group. Here, alarm sequences are extracted on a component-by-component basis. The following example illustrates the specific DTW calculation process.

[0121] 1) Get the time series of two alarms

[0122] A = [2, 5, 8] (This sequence indicates that component A generates an alarm at the 2nd, 5th, and 8th minutes)

[0123] B = [3, 6, 9] (This sequence indicates that component B generates an alarm at the 3rd, 6th, and 9th minutes)

[0124] 2) Construct a distance matrix:

[0125] For each and Calculate the absolute value of the time difference D(i,j), where i is an integer of 1, 2, or 3, and j is an integer of 1, 2, or 3. The calculation formula for D(i,j) is as follows:

[0126] ;

[0127] Table 1

[0128] A \ B 3 6 9 2 |2-3| = 1 |2-6| = 4 |2-9| = 7 5 |5-3| = 2 |5-6| = 1 |5-9| = 4 8 |8-3| = 5 |8-6| = 2 |8-9| = 1

[0129] From Table 1 we can see that each and The absolute value of the time difference. A three-row and three-column distance matrix is constructed in this way:

[0130] ;

[0131] 3) Calculate the cumulative cost matrix:

[0132] The value calculated for each location in the cumulative cost matrix is related to the value of the corresponding location in the distance matrix. The calculation is as follows:

[0133] ;

[0134] Table 2

[0135] In Table 2, the value of each position in the cumulative cost matrix is calculated based on the sum of the minimum value of the adjacent positions to its left, above, and above the left, and the value of the position in the distance matrix. For C(1,1), since there are no values for C(0,1), C(1,0), and C(0,0), these three values are considered to be 0, and C(1,1)=D(1,1). Based on the calculation of C(1,1), C(1,2), C(2,1), and C(2,2) can be further calculated, and so on, to obtain the value of each position in the cumulative cost matrix. The cumulative cost matrix is expressed as:

[0136] ;

[0137] The idea behind calculating the values at each location in the cumulative cost matrix C is to find the optimal matching path between component A and component B and calculate the cumulative time difference. This facilitates analysis of the differences in the alarm time patterns between the two components.

[0138] 4) The DTW distance W(A,B) is defined as the last element of the cumulative cost matrix C:

[0139] For the above example, the DTW distance is expressed as:

[0140] ;

[0141] Since we are analyzing the differences in the alarm time patterns between the alarm sequences of two components, we must consider the values at every position in each alarm sequence. In the cumulative cost matrix C, the last element, C(3,3), objectively reflects the DTW distance. Based on the above formula, since the last calculated value in the cumulative cost matrix is in the third row and third column, and the value in the third row and third column is 3, the DTW distance is equal to 3.

[0142] 5) Regularize the DWT distance:

[0143] During the dynamic planning process, combined with the cumulative cost matrix C, the path from the upper left corner to the lower right corner, and the backtracking path is (3,3)→(2,2)→(1,1). This process passes through a total of 3 grid points, so the number of recorded path steps L=3. The purpose of regularizing the DTW distance is to standardize the range of DTW distance values so that it is comparable with the preset distance threshold, thereby more objectively and accurately reflecting the difference in the alarm time patterns of the two components. The calculation formula for the regularized DTW distance is as follows:

[0144] ;

[0145] 6) Compare the regularized DTW distance with the preset distance threshold. If the regularized DTW distance is less than the preset distance threshold, the alarm time patterns of component A and component B are considered similar, and the causal discovery samples of component A's alarms and component B's alarms can be divided into the first group. Otherwise, if the regularized DTW distance is greater than or equal to the preset distance threshold, the alarm time patterns of component A and component B are considered significantly different, and the causal discovery samples of component A's alarms and component B's alarms can be divided into the second group.

[0146] As an example, the preset distance threshold is 5. Taking the above example where the regularized DTW=1, and comparing it with the preset distance threshold 5, it can be considered that this group of alarms belongs to the first group.

[0147] As described above, the construction of the alarm knowledge graph describes how to divide causal discovery samples into the first or second group. Multiple base learners can be used to analyze the causal discovery samples in the first group to generate a first knowledge graph, and multiple base learners can be used to analyze the causal discovery samples in the second group to generate a second knowledge graph. The reason for analyzing the causal discovery samples in the first and second groups separately to generate two knowledge graphs is that the causal discovery samples in the second group potentially contain different information than the causal discovery samples in the first group. Generating the first and second knowledge graphs separately helps to identify gaps and form a more complete alarm knowledge graph. Ultimately, the desired alarm knowledge graph is obtained by taking the union of the first and second knowledge graphs. Taking the union is itself a way to fuse and remove duplicate information from the two graphs.

[0148] The following describes the process of generating a knowledge graph by analyzing causal discovery samples using various base learners. The methods for generating the first and second knowledge graphs are essentially the same, differing only in which group the samples selected for analysis belong to. For the sake of brevity, the following example uses the generation of the first knowledge graph as an example. The description of the generation of the second knowledge graph can be found in that example.

[0149] In an optional implementation, a plurality of different base learners are used to analyze the first group of causal discovery samples to generate a first knowledge graph, including:

[0150] First, based on the alarm classification labels of the first group of causal discovery samples, a comprehensive score is performed on the candidate causal edges involved in the first group of causal discovery samples using multiple different base learners. A candidate causal edge represents a causal relationship between two alarm types to be evaluated; the two alarm types are determined based on the alarm classification labels. Next, based on the comprehensive scores of the same candidate causal edge using multiple different base learners, a comprehensive evaluation result corresponding to each candidate causal edge involved in the first group of causal discovery samples is obtained. If the comprehensive evaluation result indicates that the corresponding candidate causal edge is a causal edge, the causal edge and its associated nodes are output. Based on the output causal edges and associated nodes, a first knowledge graph is generated.

[0151] Regarding the steps for obtaining the comprehensive evaluation results of candidate causal association edges through comprehensive scoring introduced above, the specific implementation can be:

[0152] Obtain the existence consensus scores of multiple different base learners for the same candidate causal association edge; obtain the weighted aggregation scores of multiple different base learners for the same candidate causal association edge; and obtain the directional consistency scores of multiple different base learners for the candidate causal association edge. On the basis of obtaining these three types of scores, calculate the comprehensive scores of multiple different base learners for the same candidate causal association edge according to the existence consensus score, weighted aggregation score and directional consistency score. This comprehensive score is equivalent to the situation where the three different types of scores are combined. If the comprehensive score is greater than or equal to the set threshold, a comprehensive evaluation result indicating that the candidate causal association edge is a causal edge is obtained; if the comprehensive score is less than the set threshold, a comprehensive evaluation result indicating that the candidate causal association edge is a non-causal edge is obtained.

[0153] In this application, when scoring candidate causal association edges, causal discovery models such as PC (Peter Clark), GES (Greedy Equivalence Search), and LiNGAM (Linear Non-Gaussian Additive Model) can be used as base learners. The PC algorithm, GES algorithm, and LiNGAM model are classic methods in the field of causal discovery, based on constraints, scoring search, and function model assumptions, respectively, and are applicable to different data scenarios. In the actual application of the embodiments of this application, the base learners used are not limited to the PC algorithm, GES algorithm, and LiNGAM model. For ease of understanding, these three model algorithms are first briefly introduced below.

[0154] The PC algorithm is a constraint-based causal discovery algorithm that uses conditional independence tests to infer causal relationships between variables. Starting with a completely undirected graph, the algorithm gradually removes conditionally independent edges, ultimately obtaining an undirected graph (skeleton) that reflects the causal relationships between variables. Then, by identifying V-structures (i.e., three variables where one variable points to the other two without direct edges between them) and directed rules, the undirected graph is transformed into a partially directed acyclic graph (PDAG), revealing the causal direction between the variables.

[0155] The GES algorithm is a causal discovery algorithm based on score search. It greedily searches the space of graph structures to find the optimal causal graph. Starting with an empty graph, the algorithm gradually adds or removes edges to maximize a scoring function (such as the Bayesian Information Criterion (BIC). During the search, the GES algorithm considers the impact of adding and removing edges on the overall score, thereby finding the locally optimal causal graph structure.

[0156] The LiNGAM model (Linear Non-Gaussian Additive Noise Model) is a causal discovery method based on a functional model. It assumes that data is generated by a set of linear equations with non-Gaussian noise terms. The LiNGAM model estimates the coefficients of these linear equations using methods such as independent component analysis (ICA) to reveal causal relationships between variables. Due to the non-Gaussian nature of the noise terms, the LiNGAM model is able to distinguish causal directions because the statistical properties of non-Gaussian noise in the causal and anti-causal directions are different.

[0157] From the above introduction, it can be seen that the PC algorithm, GES algorithm and LiNGAM model each have their own unique principles and characteristics. In this application, the integrated learning method is used to comprehensively score the candidate causal correlation edges, with the aim of improving the credibility and reliability of the candidate causal correlation edge evaluation. For the scoring mechanism of the existence consensus score, weighted aggregation score and directional consistency score mentioned above, the performance of the same candidate causal correlation edge in different base learners is mainly considered from the three scoring dimensions. For the sake of ease of understanding, the calculation method of the existence consensus score, weighted aggregation score and directional consistency score is explained below.

[0158] (1) Existence consensus score

[0159] The definition of the existence consensus score ECS is: the frequency of occurrence of the edge in the base learner. Its formula is as follows:

[0160]

[0161] In the formula, E k It can represent the set of causal discovery samples of the first group, Represents the set E k A candidate causal edge in the dataset needs to be evaluated. K is the total number of base learners. In this embodiment, if three base learners are used, K = 3. I[ ] is an indicator function. If a base learner believes that the edge exists, the function value of this function is 1; if a base learner believes that there is no change, the function value of this function is 0. A higher value of the existence consensus score ECS indicates that more base learners agree that the edge exists, and therefore, the evaluation result is more reliable.

[0162] (2) Weighted Aggregate Scoring

[0163] The definition of weighted aggregate score WAS is: weighted average of normalized edge weights (such as confidence and effect size).

[0164] Normalization refers to mapping the weights of candidate causal edges in each base learner to the interval [0, 1]. The normalization methods for the weights of candidate causal edges in different base learners may vary. The formula for weighted aggregation to obtain the score WAS is as follows:

[0165]

[0166] In the formula, K is the total number of base learners, is the weight of the k-th base learner, represents the normalized weight of the candidate causal edge in the k-th base learner. In practical applications, It can be dynamically adjusted based on the historical performance of the base learner. In addition, the default weight value can be set: =1 / K.

[0167] (3) Directional consistency score

[0168] The definition of the Directional Consistency Score (DCS) is: the degree of consistency of edge directions in different base learners. Its formula is as follows:

[0169]

[0170] In the formula, K is the total number of base learners, Indicates that the direction of the edge is X i Pointing to X j , Indicates that the direction of the edge is X j Pointing to X i I[ ] is an indicator function. If the base learner believes that the edge exists and the direction of the edge is correct, the function value of the function is 1; if the base learner believes that the edge does not exist or exists but the direction is wrong, the function value of the function is 0.

[0171] In the above formula, if the denominator is 0, then DCS is 0. The higher the DCS, the better the base learners are at predicting the candidate causal relationship edges. There is a high consensus on the direction.

[0172] After obtaining the existence consensus score (ECS), weighted aggregation score (WAS), and directional consistency score (DCS) of a candidate causal edge, we can calculate the comprehensive score of the candidate causal edge from multiple different base learners. The following is an example implementation of calculating the comprehensive score using the formula:

[0173]

[0174] In the formula, ICS represents the comprehensive score, and β1, β2, and β3 represent the weights of the three scores: the Existential Consensus Score (ECS), the Weighted Aggregate Score (WAS), and the Directional Consistency Score (DCS). For example, β1 = 0.4, β2 = 0.4, and β3 = 0.2. These three weights can be adjusted based on the specific knowledge domain; there are no strict restrictions on their values.

[0175] Once the ICS score for a candidate causal edge has been obtained, we can then simply determine whether the edge can be considered a causal edge by comparing it to a threshold. For example, we set the threshold to 0.7. If the ICS score is greater than or equal to the threshold, the edge is retained as a causal edge; conversely, if the ICS score is less than the threshold, the edge is considered non-causal. Causal edges must be represented in the generated graph, while non-causal edges do not.

[0176] As an example, among the three base learners, PC and GES believe that edge X→Y exists, while LiNGAM believes that edge X→Y does not exist, then the calculated ECS is (1+1+0) / 3=0.67. Assume that PC and GES detect edge X→Y, and the normalized weights are 0.8 and 0.6 respectively, while LiNGAM does not detect edge X→Y, and the normalized weight is 0, based on which WAS is calculated as (0.8+0.6+0) / 3=0.47. When calculating WAS here, The value is 1 / 3. PC thinks the direction of the edge is correct, GES thinks the direction of the edge is wrong, and LiNGAM does not exist, so DCS = (1 + 0 + 0) / 2 = 0.5. Finally, calculate the X→Y comprehensive score:

[0177] ICS=0.4*0.67+0.4*0.47+0.2*0.5=0.556;

[0178] If the threshold is set to 0.5, then since 0.556 is greater than 0.5, we can conclude that the edge X→Y is a causal edge. This edge needs to be retained in the generated knowledge graph.

[0179] In practical applications, the selection of causal edges can also incorporate expert experience to screen for false-positive associations, further ensuring the accuracy and rationality of causal analysis results. In practical applications, the calculation of causal edge weights can be determined using Naive Bayesian theory. Finally, a first knowledge graph is generated based on the entity nodes representing each alarm type, causal edges, and causal edge weights. Based on a similar implementation, a second set of causal discovery samples is gradually scored to generate a second knowledge graph. The first and second knowledge graphs are then unioned to obtain the alarm knowledge graph. The generated alarm knowledge graph can be stored in a graph database, such as the Nebula database. Figure 3A schematic diagram of the construction process of an alarm knowledge graph based on ensemble learning provided in this application embodiment, combined with Figure 3 From top to bottom, we can see the entire process: constructing causal discovery samples, analyzing and scoring multiple base learners, integrating the evaluation results of multiple base learners, and finally combining expert experience to output causal nodes and causal edges. Finally, through the calculation of causal edge weights, the alert knowledge graph is constructed and stored in the Nebula database. It should be noted that a causal node is the physical node connecting the two ends of a causal edge, representing a specific alert type.

[0180] In this embodiment, causal discovery samples are grouped during the initial preparation phase of building the alert knowledge graph. The first group (predicted correlation grouping) reduces false positives, while the second group (random grouping) expands the scope of explorable causal relationships. The combination of these two ensures high confidence while uncovering unknown causal relationships. Through an ensemble learning scoring mechanism, expert false positive screening, and naive Bayesian weighting of causal edges, the accuracy of the data ultimately stored in the alert knowledge graph is improved.

[0181] In optional implementations, methods for locating the root cause of alarms in operation and maintenance scenarios also include:

[0182] Obtain the software and hardware knowledge graph for operation and maintenance scenarios;

[0183] For new alarm data, the root cause of the alarm is located using the corrected alarm knowledge graph. Specifically, the corrected alarm knowledge graph and the obtained software and hardware knowledge graph can be combined to locate the root cause of the alarm for the new alarm data.

[0184] To facilitate understanding of the combined application of the alarm knowledge graph and the hardware and software knowledge graph, the following explanation is provided. When a real-time alarm is triggered, it is initially classified using the alarm classification model and assigned to an abstract business system. Within a set time window, the hardware and software knowledge graphs are combined with the alarm knowledge graph to perform relationship mining on the alarm propagation path and impact range, automatically discovering alarm correlations and locating the root cause.

[0185] Combined with the scene example:

[0186] Scenario description: A large number of packet loss alarms appear in the switch logs of a data center, and the network throughput of multiple servers drops abnormally.

[0187] Analysis process:

[0188] (1) Alarm classification: Based on historical data, the packet loss alarm is classified as a network device packet loss failure, and the server network throughput abnormality is classified as a service delay alarm.

[0189] (2) Merge abstract business systems: Determine whether the affected servers belong to the same tenant (e.g., cloud database service).

[0190] (3) Transmission path analysis:

[0191] (3.1) Based on the hardware and software knowledge graph, we determine that the relevant servers are connected to the same core switch. This means that there is a link in the hardware and software knowledge graph: switch -> server. Both the switch and the server are physical nodes in the hardware and software knowledge graph.

[0192] (3.2) Query the alarm knowledge graph to find the link: Network device packet loss fault -> Service delay alarm. Among them, Network device packet loss fault and Service delay alarm are two entity nodes representing alarm types in the alarm knowledge graph.

[0193] (3.3) Considering the results of (3.1) and (3.2), we found that the large number of packet loss alarms in the switch logs and the abnormal server network throughput have a propagation relationship in both the hardware and software knowledge graphs and the alarm knowledge graphs. Therefore, we believe that there is a causal relationship between the large number of packet loss alarms in the switch logs and the abnormal server network throughput.

[0194] (4) Root cause location: It was eventually discovered that the core switch caused data packet loss due to traffic anomalies, which in turn affected the network throughput of the business server.

[0195] Figure 4 A flowchart of locating the root cause of an alarm provided in an embodiment of the present application. Figure 4 As can be seen in the figure, after the real-time alarm arrives, it is sliced according to the time and judged whether it belongs to the existing classification. If so, it is classified into the existing classification, that is, Figure 4 This refers to "alarm convergence" in [1]. If the alarm does not fall into an existing category, a new alarm category must be created. After completing the specific classification, the alarms are correlated and the root cause is located by combining the software and hardware knowledge graph with the alarm knowledge graph, ultimately presenting the located root cause. For example, in the scenario above, where "core switches experience packet loss due to traffic anomalies, which in turn affects the network throughput of business servers," the root cause can be highlighted in a specific form of text or image, thereby presenting the root cause. Figure 4 Nebula is used as a database to store software and hardware knowledge graphs and alarm knowledge graphs. Figure 4 The alarm knowledge graph in refers to the alarm knowledge graph that has been subjected to graph correction according to the embodiment described above.

[0196] It should be noted that in the alarm root cause location method in the operation and maintenance scenario provided by the embodiment of the present application, the real-time update of the knowledge graph can also be achieved through the logic of dynamic incremental learning. In addition, the embodiment of the present application also provides a technical solution for offline correction of the knowledge graph through the expert experience interface, as well as a technical solution for online correction of the knowledge graph based on the real-time evaluation of the alarm root cause. Among them, the online correction of the knowledge graph based on the implementation evaluation of the alarm root cause is mainly achieved by adjusting the sigmoid reputation scoring function, which can be referred to Figure 2A The three adjustment ideas shown are as follows. This application can complete the offline and online hybrid verification of the knowledge graph and realize the dynamic incremental learning and update of the knowledge graph. Figure 5 A schematic diagram of an embodiment of the present application providing a method for updating a knowledge graph using an offline-online hybrid correction mechanism to locate the root cause of an alarm. Figure 5 The process in the leftmost column reflects the technical implementation of atlas visualization and offline correction based on expert experience; Figure 5 The middle column in the figure reflects the dynamic graph update, which is mainly completed through incremental learning based on the existing graph. Figure 5 The right column reflects the accuracy of the real-time alarm root cause location, completes the online feedback of the root cause verification and adjusts the reputation score, and ensures the accuracy and reliability of the alarm knowledge graph through online correction.

[0197] Below, we first explain how to implement dynamic incremental learning updates of knowledge graphs.

[0198] The inventors believe that the dynamic incremental learning logic of the knowledge graph is a key mechanism to ensure that the knowledge graph remains up to date in an ever-changing environment. As a graph structure used to represent entities and their relationships, with the continuous updating and expansion of data, static knowledge graphs are difficult to cope with the rapid emergence of new knowledge and the obsolescence of old knowledge. Therefore, the introduction of dynamic incremental learning logic is particularly important. The dynamic incremental learning proposed in the technical solution of this application is a method for gradually learning and updating the knowledge graph. It can update the graph in time when new data is received without rebuilding the entire graph from scratch. This not only improves efficiency, but also ensures that the knowledge graph can reflect the latest information and relationships in a timely manner. The main steps of dynamic incremental learning include data cleaning, relationship mining, data fusion and graph updating.

[0199] First, data cleaning filters out noisy data and erroneous information to ensure the quality of input data. Next, relationship mining uses causal inference algorithms to extract alarm causal relationships from newly added alarm data. Data fusion integrates the newly extracted alarm causal relationships with the existing knowledge graph. Graph updates formally add new data to the knowledge graph. Dynamic incremental learning logic not only processes large amounts of new data in a short period of time, but also improves the overall quality and coverage of the knowledge graph through continuous updates and optimization. For example, as new resources and alarms are managed, the alarm knowledge graph can be promptly updated to reflect the latest alarm relationships.

[0200] Next, the technical implementation of offline correction of knowledge graph based on expert experience in this application is explained.

[0201] The expert experience interface plays a vital role in the offline correction of the knowledge graph. Although the AI-based automatic graph update method has played a major role in the construction of the knowledge graph, this update method still has certain limitations, especially when dealing with complex alarm scenarios and information with strong scenario-based information. At this time, the knowledge and experience of human experts can greatly improve the accuracy and quality of the knowledge graph. Offline correction is a method of correcting and optimizing the knowledge graph in a non-real-time environment. The expert experience interface provides an intuitive and interactive platform for operation and maintenance experts, enabling them to mark and correct errors and inaccuracies in the knowledge graph, and enter missing relationships. The expert experience interface mainly presents the following functional modules:

[0202] (1) Data browsing and querying: Experts can browse and query the data in the knowledge graph through the interface. This function enables experts to fully understand the structure, content, and relationships of the graph, thereby discovering potential problems and improvement points.

[0203] (2) Error marking and correction: When experts find errors in the knowledge graph, they can mark and correct them directly on the interface.

[0204] (3) New knowledge and relationships: Experts can also add new entities and relationships to the knowledge graph through the interface. This feature is particularly suitable for introducing new discoveries and new knowledge. For example, as new scientific research results emerge, experts can add this new knowledge to the graph to keep it up-to-date and accurate.

[0205] (4) Version control and history recording: The expert experience interface includes version control and history recording functions. Each modification is recorded and a new version is generated. This not only makes it easier to track and review the modification history, but also allows you to restore to the previous version when necessary, ensuring the stability and reliability of the graph.

[0206] (5) Collaboration and review: To ensure the quality and consistency of corrections, the expert experience interface usually supports collaboration and review functions. Multiple experts can work together to calibrate and optimize the same part of the map. Each modification can be reviewed and confirmed by other experts to ensure the accuracy and rationality of the modification.

[0207] Offline correction of data in the knowledge graph through the expert experience interface not only fully leverages the knowledge and experience of human experts, but also allows for in-depth verification and modification without impacting the system's real-time performance. This process helps to offset the shortcomings of automated algorithms and ensure that the knowledge graph maintains high quality even in complex and unique situations. Furthermore, the expert experience interface serves as a platform for accumulating and transferring expert knowledge during the knowledge graph construction process. The annotations and corrections made by experts during the correction process can serve as valuable knowledge resources for subsequent graph construction and optimization. This knowledge accumulation and transfer mechanism helps to enhance the professional level and work efficiency of the entire team.

[0208] Figure 6 An implementation architecture diagram for locating the root cause of an alarm using a multi-dimensional knowledge graph, provided in an embodiment of the present application. Figure 6 The project demonstrates the construction, correction, and application of graphs from the perspectives of data layer, graph construction, real-time analysis, graph storage, and adjustment. CMDB, short for Configuration Management Database, is a logical database used to store and manage various configuration information for devices within an enterprise IT architecture. It includes physical relationships, real-time communication relationships, non-real-time communication relationships, and dependencies between various configurations. Figure 6 The arrow from the CMDB data to the software and hardware knowledge graph indicates that the software and hardware knowledge graph used in this application can be constructed based on the information and relationships in the CMDB data. For the use of historical alarm data and real-time alarms, please refer to Figure 3 and Figure 4 For more information on root cause assessment-based atlas correction, please refer to Figure 2A and Figure 5 The rightmost column in the , will not be described here.

[0209] In O&M scenarios, existing technologies generally rely on single system topology data or local alarm information, making it difficult to leverage both the actual hardware and software topology and abstract correlation information. This makes it difficult to fully reflect system status in complex scenarios. Traditional methods often rely on O&M personnel to make judgments based on historical experience, making it difficult to effectively leverage large-scale historical data for automatic learning and knowledge mining, thereby reducing the accuracy of root cause identification and response speed. During system operation, alarm data and relationships are constantly changing, making it difficult for static knowledge graphs to reflect the latest developments in real time. Existing technologies lack a hybrid correction mechanism that can leverage both offline expert correction and online updates based on real-time evaluation.

[0210] In this application, through the construction of a multi-dimensional knowledge graph, the physical topology of the system and the abstract alarm data are fully integrated, which can more accurately reflect the causal relationship between alarms, thereby realizing automatic root cause location and improving the accuracy of root cause location. A dynamic incremental learning mechanism and an offline-online hybrid correction method are adopted to ensure that the knowledge graph can be updated in a timely manner, adapt to system environment and data changes, and reduce information lag. The reputation scoring system design based on the Sigmoid function realizes the rapid punishment of the wrong root cause and the gradual reward of the correct root cause, effectively preventing the accumulation and spread of erroneous information in the knowledge graph. By integrating machine learning, causal analysis and expert correction methods, the dependence on pure manual experience is reduced, the intelligence of alarm analysis is realized, and the speed and efficiency of fault response are improved.

[0211] Based on the method for locating the root cause of an alarm in an operation and maintenance scenario introduced in the above embodiment, the present application also provides a device for locating the root cause of an alarm in an operation and maintenance scenario. Figure 7 is a structural diagram of the device, as shown in Figure 7 As shown, the device includes:

[0212] A graph acquisition module 701 is used to acquire an alarm knowledge graph; the alarm knowledge graph is constructed based on an operation and maintenance scenario, the entities in the alarm knowledge graph represent alarm types, and the edges between entities represent the link relationships between alarm types;

[0213] The data acquisition module 702 is used to obtain error statistics of the alarm knowledge graph; the error statistics include edges of each root cause of the positioning error within the statistical period and corresponding error frequency data;

[0214] The score acquisition module 703 is used to obtain the reputation score of each edge in the alarm knowledge graph that locates the root cause of the error; the reputation score is positively correlated with the accuracy of the corresponding edge in locating the root cause;

[0215] An adjustment factor determination module 704 is configured to determine a dynamic sensitivity adjustment factor of a corresponding edge based on the error frequency data;

[0216] A function adjustment module 705 is configured to adjust the reputation scoring function of the corresponding edge using the dynamic sensitivity adjustment factor to obtain an adjusted reputation scoring function corresponding to each edge of the root cause of the positioning error;

[0217] A graph correction module 706 is configured to correct the reputation score of each edge of the root cause of the positioning error using the corresponding adjusted reputation scoring function to obtain a corrected reputation score for each edge of the root cause of the positioning error, so as to correct the alarm knowledge graph;

[0218] The root cause location module 707 is used to locate the root cause of the alarm using the corrected alarm knowledge graph for the new alarm data.

[0219] In an optional implementation, the alarm root cause location device in the operation and maintenance scenario further includes:

[0220] A correction time interval acquisition module is used to obtain the time interval between the most recent correction time and the current time of the reputation score of each edge of the root cause of the positioning error;

[0221] A time decay factor acquisition module, configured to obtain a time decay factor of each edge of a root cause of a positioning error according to each of the time intervals;

[0222] The spectrum correction module 706 is specifically used to:

[0223] For each edge of the root cause of the positioning error, the product of the corresponding time decay factor and the reputation score is obtained, and based on the product, the reputation score is corrected using the corresponding adjusted reputation scoring function.

[0224] In an optional implementation, the alarm root cause location device in the operation and maintenance scenario further includes:

[0225] An error level identification module is used to identify the error level of each edge that is the root cause of the positioning error;

[0226] a first increment determination module, configured to determine, for edges of a first error level, an increment of an independent variable of an adjusted reputation scoring function by taking a logarithm of the severity of the error;

[0227] The second increment determination module is used to determine the independent variable increment of the adjusted reputation scoring function by multiplying the severity of the error by itself for the edge of the second error level; wherein the error degree of the first error level is higher than the error degree of the second error level.

[0228] In an optional implementation, the alarm root cause location device in the operation and maintenance scenario further includes:

[0229] a third increment determination module, configured to determine an increment of an independent variable of the reputation scoring function by taking the square root of the correct contribution of each correctly located root cause edge within the statistical period;

[0230] The score acquisition module 703 is further used to obtain the reputation score of each edge of the correct root cause in the alarm knowledge graph;

[0231] The graph correction module 706 is further configured to adjust the reputation scoring function of each edge with a correctly located root cause within the statistical period using the initial sensitivity adjustment factor to correct the reputation score of the corresponding edge.

[0232] In an optional implementation, the apparatus further includes a graph construction module 708, which includes:

[0233] A sample construction unit, used to construct historical alarm data into multiple causal discovery samples according to time;

[0234] a sample grouping unit, configured to group the causal discovery samples based on the correlation characteristics of the causal discovery samples; wherein the causal discovery samples whose correlation characteristics meet a preset condition are grouped into a first group, and the causal discovery samples that do not meet the preset condition are grouped into a second group;

[0235] An integrated learning unit is configured to analyze the first group of causal discovery samples using a plurality of different base learners to generate a first knowledge graph; and to analyze the second group of causal discovery samples using the plurality of different base learners to generate a second knowledge graph; in the first knowledge graph and the second knowledge graph, entities represent alarm types, and edges between entities represent link relationships between alarm types;

[0236] A graph processing unit is used to obtain the union of the first knowledge graph and the second knowledge graph to obtain the alarm knowledge graph.

[0237] In an optional implementation, the integrated learning unit is specifically configured to:

[0238] Based on the alarm classification labels of the first group of causal discovery samples, using the multiple different base learners to comprehensively score candidate causal association edges involved in the first group of causal discovery samples; the candidate causal association edges represent a causal relationship to be evaluated between two alarm types; the two alarm types are determined based on the alarm classification labels;

[0239] Obtaining a corresponding comprehensive evaluation result of each candidate causal association edge involved in the causal discovery samples of the first group based on the comprehensive scores of the same candidate causal association edge by the multiple different base learners;

[0240] If the comprehensive evaluation result indicates that the corresponding candidate causal association edge is a causal edge, outputting the causal edge and the associated nodes of the causal edge;

[0241] Based on the output causal edges and associated nodes, the first knowledge graph is generated.

[0242] In an optional implementation, the integrated learning unit is specifically configured to:

[0243] Obtaining existence consensus scores of the multiple different base learners for the same candidate causal association edge; obtaining weighted aggregate scores of the multiple different base learners for the same candidate causal association edge; and obtaining directional consistency scores of the multiple different base learners for the candidate causal association edge;

[0244] Calculating, based on the existence consensus score, the weighted aggregation score, and the direction consistency score, a comprehensive score of the multiple different base learners for the same candidate causal association edge;

[0245] If the comprehensive score is greater than or equal to the set threshold, a comprehensive evaluation result indicating that the candidate causal association edge is a causal edge is obtained; if the comprehensive score is less than the set threshold, a comprehensive evaluation result indicating that the candidate causal association edge is a non-causal edge is obtained.

[0246] In an optional implementation, the preset conditions include: there is a dependency relationship between services, and / or the regularized dynamic time warping distance is less than a preset distance threshold.

[0247] In an optional implementation, in the alarm root cause location device in the operation and maintenance scenario:

[0248] The graph acquisition module 701 is also used to obtain the software and hardware knowledge graph of the operation and maintenance scenario;

[0249] The root cause location module 707 is specifically configured to:

[0250] The corrected alarm knowledge graph and the software and hardware knowledge graph are used in combination to locate the root cause of the alarm based on the new alarm data.

[0251] Based on the aforementioned method embodiments and apparatus embodiments, the present application further provides an alarm root cause location device in an operation and maintenance scenario, the device comprising a processor and a memory communicatively connected to each other;

[0252] The memory stores a computer program;

[0253] The processor is used to run the computer program to implement the alarm root cause location method in the operation and maintenance scenario as introduced in the method embodiment.

[0254] It should be noted that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and equipment embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments. The device and equipment embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components indicated as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without expending creative work.

[0255] The above is merely one specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A method for locating the root cause of an alarm in an operation and maintenance scenario, characterized in that: include: Obtain the alarm knowledge graph; The alarm knowledge graph is constructed based on operation and maintenance scenarios. The entities in the alarm knowledge graph represent alarm types, and the edges between entities represent the link relationships between alarm types. Obtain error statistics of the alarm knowledge graph; the error statistics include edges of each root cause of positioning errors within a statistical period and corresponding error frequency data; Obtaining a reputation score for each edge in the alarm knowledge graph that locates the root cause of the error; wherein the reputation score is positively correlated with the accuracy of the corresponding edge in locating the root cause; determining a dynamic sensitivity adjustment factor of a corresponding edge based on the error frequency data; Using the dynamic sensitivity adjustment factor, adjust the reputation scoring function of the corresponding edge to obtain an adjusted reputation scoring function corresponding to each edge of the root cause of the positioning error; For each edge of the root cause of the positioning error, correct the reputation score using the corresponding adjusted reputation scoring function to obtain a corrected reputation score for the edge of each root cause of the positioning error, so as to correct the alarm knowledge graph; For new alarm data, the root cause of the alarm is located using the corrected alarm knowledge graph.

2. The method according to claim 1, characterized in that Before correcting the reputation score of each edge of the root cause of the positioning error using the corresponding adjusted reputation scoring function, the method further includes: Get the time interval between the last correction time and the current time of the reputation score of each edge that locates the root cause of the error; Obtaining a time attenuation factor of each edge of the root cause of the positioning error according to each of the time intervals; For each edge of the root cause of the positioning error, the reputation score is corrected using the corresponding adjusted reputation scoring function, including: For each edge of the root cause of the positioning error, the product of the corresponding time decay factor and the reputation score is obtained, and based on the product, the reputation score is corrected using the corresponding adjusted reputation scoring function.

3. The method according to claim 1, characterized in that Before correcting the reputation score of each edge of the root cause of the positioning error using the corresponding adjusted reputation scoring function, the method further includes: Identify the error level of each edge that is the root cause of the positioning error; For the edges of the first error level, determine the increment of the independent variable of the adjusted reputation scoring function by taking the logarithm of the severity of the error; For the edge of the second error level, the independent variable increment of the adjusted reputation scoring function is determined by multiplying the severity of the error by itself; wherein the error degree of the first error level is higher than the error degree of the second error level.

4. The method according to claim 3, characterized in that Also includes: For each edge with the correct root cause located within the statistical period, determine the independent variable increment of the reputation scoring function by taking the square root of the correct contribution; Obtaining the reputation score of each edge in the alarm knowledge graph that locates the correct root cause; For each edge whose root cause is correctly located within the statistical period, the reputation scoring function of the corresponding edge is adjusted using the initial sensitivity adjustment factor to correct the reputation score of the corresponding edge.

5. The method according to claim 1, wherein The steps of constructing the alarm knowledge graph include: Construct historical alarm data into multiple causal discovery samples based on time; Based on the correlation characteristics of the causal discovery samples, the causal discovery samples are grouped; wherein the causal discovery samples whose correlation characteristics meet the preset conditions are divided into the first group, and the causal discovery samples that do not meet the preset conditions are divided into the second group; Using multiple different base learners to analyze the first group of causal discovery samples to generate a first knowledge graph; and using the multiple different base learners to analyze the second group of causal discovery samples to generate a second knowledge graph; in the first knowledge graph and the second knowledge graph, entities represent alarm types, and edges between entities represent link relationships between alarm types; The first knowledge graph and the second knowledge graph are combined to obtain the alarm knowledge graph.

6. The method according to claim 5, characterized in that The method of using a plurality of different base learners to analyze the first group of causal discovery samples to generate a first knowledge graph includes: Based on the alarm classification labels of the first group of causal discovery samples, using the multiple different base learners to comprehensively score candidate causal association edges involved in the first group of causal discovery samples; the candidate causal association edges represent a causal relationship to be evaluated between two alarm types; the two alarm types are determined based on the alarm classification labels; Obtaining a corresponding comprehensive evaluation result of each candidate causal association edge involved in the causal discovery samples of the first group based on the comprehensive scores of the same candidate causal association edge by the multiple different base learners; If the comprehensive evaluation result indicates that the corresponding candidate causal association edge is a causal edge, outputting the causal edge and the associated nodes of the causal edge; Based on the output causal edges and associated nodes, the first knowledge graph is generated.

7. The method according to claim 6, characterized in that The comprehensive evaluation results corresponding to the candidate causal association edges involved in the causal discovery samples of the first group are obtained based on the comprehensive scores of the same candidate causal association edge by the multiple different base learners, including: Obtaining existence consensus scores of the multiple different base learners for the same candidate causal association edge; obtaining weighted aggregate scores of the multiple different base learners for the same candidate causal association edge; and obtaining directional consistency scores of the multiple different base learners for the candidate causal association edge; Calculating, based on the existence consensus score, the weighted aggregation score, and the direction consistency score, a comprehensive score of the multiple different base learners for the same candidate causal association edge; If the comprehensive score is greater than or equal to the set threshold, a comprehensive evaluation result indicating that the candidate causal association edge is a causal edge is obtained; if the comprehensive score is less than the set threshold, a comprehensive evaluation result indicating that the candidate causal association edge is a non-causal edge is obtained.

8. The method according to any one of claims 5 to 7, characterized in that The preset conditions include: there is a dependency relationship between the services, and / or the regularized dynamic time warping distance is less than a preset distance threshold.

9. The method according to any one of claims 1 to 7, characterized in that The method further comprises: Obtaining a software and hardware knowledge graph for the operation and maintenance scenario; The method of locating the root cause of the alarm using the corrected alarm knowledge graph for the new alarm data includes: The corrected alarm knowledge graph and the software and hardware knowledge graph are used in combination to locate the root cause of the alarm based on the new alarm data.

10. A device for locating the root cause of an alarm in an operation and maintenance scenario, characterized in that: include: Graph acquisition module, used to obtain alarm knowledge graph; The alarm knowledge graph is constructed based on operation and maintenance scenarios. The entities in the alarm knowledge graph represent alarm types, and the edges between entities represent the link relationships between alarm types. A data acquisition module is used to obtain error statistics of the alarm knowledge graph; the error statistics include edges of each root cause of positioning errors within a statistical period and corresponding error frequency data; A score acquisition module is used to obtain the reputation score of each edge in the alarm knowledge graph that locates the root cause of the error; the reputation score is positively correlated with the accuracy of the corresponding edge in locating the root cause; an adjustment factor determination module, configured to determine a dynamic sensitivity adjustment factor of a corresponding edge based on the error frequency data; A function adjustment module, configured to adjust the reputation scoring function of the corresponding edge using the dynamic sensitivity adjustment factor to obtain an adjusted reputation scoring function corresponding to each edge of the root cause of the positioning error; A graph correction module, configured to correct the reputation score of each edge of each root cause of the positioning error using the corresponding adjusted reputation scoring function to obtain a corrected reputation score of each edge of the root cause of the positioning error, so as to correct the alarm knowledge graph; The root cause location module is used to locate the root cause of alarms based on new alarm data using the corrected alarm knowledge graph.

11. A device for locating the root cause of an alarm in an operation and maintenance scenario, characterized in that: include: a processor and memory communicatively connected to each other; The memory stores a computer program; The processor is configured to run the computer program to implement the alarm root cause locating method in the operation and maintenance scenario according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Real-time root cause analysis method based on operation and maintenance knowledge graph

    CN116225760A

  • Alarm storm-oriented root cause positioning method, system and device and medium

    CN117880069A

  • System and method for determining and managing reputation of entities and industries

    US20220261825A1

  • Data set and algorithm validation, bias characterization, and valuation

    US20250088542A1