A root cause positioning-oriented evidence screening method and related device
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-08-11
AI Technical Summary
多数方案仅依据单条证据自身的异常分值进行筛选,以至于最终纳入结果的证据中常混杂冗余证据、噪声证据和冲突证据,影响根因定位的准确性和可解释性
本申请通过基于多源数据确定候选证据集合,并综合考虑候选证据的解释能力、候选证据之间的冲突程度以及候选证据集合对根因结论的稳定支撑能力,对候选证据进行筛选,能够避免仅依据单条证据异常分值进行筛选所导致的冗余证据、噪声证据和冲突证据混入问题,从而提高筛选得到的初始证据集合对目标异常事件的解释质量。进一步地,本申请通过对初始证据集合中的候选证据逐条构造反事实证据集合,并结合回放判断反事实证据集合是否仍满足预设稳定性要求和预设解释能力要求,从而删除对根因结论稳定成立并非必要的候选证据,得到目标证据集合。由此,分阶段地筛选出稳定支撑根因结论的核心证据,提升根因定位结果的准确性、可解释性和可复现性,并有助于降低冗余证据对运维分析和复盘过程的干扰。
Smart Images

Figure CN122548577A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent operation and maintenance technology, and in particular to an evidence screening method and related apparatus for root cause localization. Background Technology
[0002] With the widespread adoption of cloud-native architectures in microservices, containers, service meshes, and computing infrastructure, cloud-native systems continuously generate multi-source data during operation, including metrics, logs, call chains, configuration changes, and release events. When system anomalies occur, operations and maintenance platforms typically need to conduct root cause analysis based on this multi-source data to help operations and maintenance personnel quickly identify the source of the fault, analyze the propagation path, and complete fault handling.
[0003] Existing root cause localization techniques typically extract evidence from multi-source data, but they still have significant shortcomings in evidence screening. Specifically, existing techniques lack a unified quantitative evaluation mechanism for evidence quality. Most schemes rely solely on the outlier scores of individual pieces of evidence for screening, resulting in a mixture of redundant, noisy, and conflicting evidence in the final results, affecting the accuracy and interpretability of root cause localization. Summary of the Invention
[0004] To address the aforementioned issues, this application provides an evidence screening method and related apparatus for root cause localization, aiming to screen out stable evidence supporting root cause conclusions and improve the accuracy and interpretability of root cause localization.
[0005] The embodiments of this application disclose the following technical solutions: The first aspect of this application provides a root cause localization-oriented evidence screening method, the method comprising: Acquire multi-source data related to the target anomalous event, and determine a set of candidate evidence based on the multi-source data; Based on the explanatory power of the candidate evidence in the candidate evidence set for the target anomalous event, the degree of conflict between the candidate evidence, and the stable support of the candidate evidence set for the root cause conclusion, evidence is screened from the candidate evidence set to obtain an initial evidence set. For each candidate piece of evidence in the initial evidence set, a counterfactual evidence set is constructed by deleting the corresponding candidate piece of evidence, and the counterfactual evidence set is determined by replay to see if it meets the preset stability requirements and preset explanatory power requirements. If the counterfactual evidence set satisfies the preset stability requirement and the preset explanatory power requirement, the corresponding candidate evidence is deleted from the initial evidence set. Repeat the candidate evidence deletion judgment until there is no candidate evidence that still meets the preset stability requirement and the preset explanatory power requirement after deletion, and obtain the target evidence set.
[0006] A second aspect of this application provides an evidence screening device for root cause localization, the device comprising: The evidence set acquisition module is used to acquire multi-source data related to the target abnormal event and determine a candidate evidence set based on the multi-source data. The first screening module is used to screen evidence from the candidate evidence set based on the explanatory power of the candidate evidence in the candidate evidence set for the target abnormal event, the degree of conflict between the candidate evidence, and the stable support power of the candidate evidence set for the root cause conclusion, to obtain an initial evidence set. The counterfactual pruning module is used to construct a counterfactual evidence set after deleting the corresponding candidate evidence for each candidate evidence in the initial evidence set, and to determine whether the counterfactual evidence set meets preset stability requirements and preset explanatory power requirements by replaying the evidence. If the counterfactual evidence set meets the preset stability requirements and preset explanatory power requirements, the corresponding candidate evidence is deleted from the initial evidence set. The candidate evidence deletion judgment is repeated until there is no candidate evidence that still meets the preset stability requirements and preset explanatory power requirements after deletion, thus obtaining the target evidence set.
[0007] Compared with the prior art, this application has the following beneficial effects: This application determines a candidate evidence set based on multi-source data and comprehensively considers the explanatory power, the degree of conflict between candidate evidence, and the stable support of the candidate evidence set for the root cause conclusion. This screening process avoids the problem of redundant, noisy, and conflicting evidence being mixed in, which can occur when screening based solely on the anomaly score of a single piece of evidence. This improves the explanatory quality of the initial evidence set for the target anomaly event. Furthermore, this application constructs a counterfactual evidence set for each candidate piece of evidence in the initial evidence set and uses playback to determine whether the counterfactual evidence set still meets the preset stability and explanatory power requirements. This process removes candidate evidence that is not essential for the stable validity of the root cause conclusion, resulting in the target evidence set. Thus, core evidence that stably supports the root cause conclusion is screened in stages, improving the accuracy, interpretability, and reproducibility of the root cause localization results, and helping to reduce the interference of redundant evidence on the operation and maintenance analysis and review process. Attached Figure Description
[0008] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0009] Figure 1 A flowchart illustrating an evidence screening method for root cause localization provided in this application embodiment; Figure 2 A flowchart illustrating an example method for determining a set of candidate evidence provided in this application embodiment; Figure 3 A flowchart for constructing an evidence chain based on a target set of evidence is provided in an embodiment of this application; Figure 4 This is a schematic diagram of an evidence screening device for root cause localization provided in an embodiment of this application. Detailed Implementation
[0010] The evidence screening method and related apparatus for root cause localization provided in this application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the following implementation is used to illustrate the technical solution of this application, and not to limit the scope of protection of this application. Without departing from the technical concept of this application, those skilled in the art can make equivalent substitutions or adaptive adjustments to the relevant step sequence, parameter form, evaluation index and module division.
[0011] This application addresses anomaly analysis in cloud-native systems. During system operation, it collects multi-source data, including metrics, logs, call chains, configuration changes, and published events, and constructs a candidate evidence set around the target anomaly event. Instead of directly filtering based solely on the anomaly score of a single piece of evidence, it performs phased screening of candidate evidence, considering the explanatory power of the candidate evidence for the target anomaly event, the degree of conflict between candidate evidence, and the stable support of the candidate evidence set for the root cause conclusion. This process yields a target evidence set that stably supports the root cause conclusion. The target evidence set can be considered as the minimum stable evidence set obtained by implementing the technical solution of this application.
[0012] Furthermore, some implementations of the technical solution in this application can construct a minimal chain of evidence by combining topological relationships and time information after obtaining the target evidence set, thereby assisting operations and maintenance personnel in understanding the anomaly propagation process and conducting retrospective analysis. Based on the above technical solution, this application can reduce the interference of redundant evidence, noisy evidence, and conflicting evidence on root cause localization results, and improve the accuracy, interpretability, and reproducibility of root cause localization.
[0013] The following description uses method embodiments. The process of the evidence screening method for root cause localization described in the embodiments of this application can be divided into steps S101 to S105, see below. Figure 1The root cause localization-oriented evidence screening method can be executed by a server, an operation and maintenance analysis engine in a cloud platform, a root cause localization system, an evidence screening device, or a processor deployed in an electronic device. Accordingly, the steps in the root cause localization-oriented evidence screening method described below can also be executed by the corresponding modules in the device embodiment.
[0014] Step S101: Obtain multi-source data related to the target abnormal event, and determine a set of candidate evidence based on the multi-source data.
[0015] In one implementation of step S101, multi-source data related to anomaly analysis can be obtained first, and candidate evidence with spatiotemporal and semantic correlation with the current target anomaly can be extracted from the multi-source data and collected into a candidate evidence set.
[0016] The target abnormal event can be an alarm event, performance degradation event, service unavailability event, resource abnormal event, or other abnormal event that requires root cause localization in the cloud-native system.
[0017] Multi-source data can include, but is not limited to: metric data, log data, call chain data, link topology data, configuration change data, release event data, work order event data, resource status data, etc.
[0018] Step S101 extracts candidate evidence with strong correlation to the target anomalous event from the original multi-source data, reducing irrelevant and low-quality data that needs to be processed in subsequent screening processes, and improving overall solution efficiency and screening quality. The resulting set of candidate evidence provides a basis for subsequent initial evidence screening based on interpretability, conflict level, and stability support capability.
[0019] In an optional implementation of step S101, the observation time window can be determined based on the alarm time of the target anomaly. For example, taking the alarm trigger time as the center, a first preset time length is extended forward and a second preset time length is extended backward to obtain the time range for observing the occurrence and propagation process of the anomaly. Then, evidence corresponding to data whose time range intersects with the observation time window is extracted from multi-source data to form a time-based preliminary evidence set. This process can eliminate evidence that is obviously unrelated to the anomaly occurrence time in advance, thereby reducing subsequent computational overhead and improving the relevance of the candidate evidence set. A more detailed explanation of this will follow.
[0020] Let the alarm time for the target abnormal event be... The observation time window can be represented as:
[0021] in, Indicates the first preset time length. This indicates the second preset time length.
[0022] In the optional implementation, this application first selects a time interval (in terms of time range) from the multi-source data. Starting with (To terminate) falling within the observation time window All the evidence within it, resulting in a time-based initial screening evidence set. :
[0023] Among the symbols This indicates that the time interval of the evidence and the observation time window have a non-empty intersection.
[0024] In an optional implementation of step S101, based on the aforementioned initial time screening, the alarm context node set corresponding to the anomaly event can be further determined by combining the alarm entity set corresponding to the target anomaly event and the service and resource topology diagram.
[0025] The alert entity set can include: service instances that trigger alerts, containers, hosts, network nodes, middleware components, or database nodes, etc. The service and resource topology graph can depict the dependencies, invocation relationships, deployment relationships, or propagation relationships between different services, components, and resource entities. Based on the alert entity set and the service and resource topology graph, a set of nodes with propagation associations to the target abnormal event can be extracted within a preset topology distance range as the alert context node set.
[0026] Subsequently, evidence from the initial time-based evidence set that the source entity is located within the alarm context node set is selected to form a context-dependent evidence set. By introducing topological context constraints, evidence that is temporally close but clearly unrelated in propagation can be further excluded, improving the specificity of the evidence set for anomaly localization tasks. A more detailed explanation of this point follows.
[0027] Let the set of service or resource entities directly involved in this alarm (i.e., the alarm entity set) be . The service and resource topology of a cloud-native system is G = (V, E). This application can determine the alarm context node set through topological neighborhood expansion. In one implementation, the alarm context node set can be defined. for:
[0028] in, This represents the set of nodes from node v to alarm nodes in graph G. The shortest path length; This is the preset maximum topology distance, used to limit the scope of context expansion.
[0029] Therefore, based on the alarm context node set, an evidence set related to the alarm context topology can be obtained. This requires that the entity from which such evidence originates must be located in the alarm context node set. Within. That is:
[0030] Combining temporal screening and topological correlation, through the aforementioned temporal screening evidence set... and the set of evidence related to the alarm context topology Finding the intersection will yield the set of context-related evidence. The format is as follows:
[0031] In one implementation, the context-related evidence set can be directly used as the candidate evidence set.
[0032] In addition, some alternative implementations can pre-prune the context-related evidence set to remove evidence that does not meet the requirements for explanatory power or data source quality, as well as duplicate evidence. Explanatory power here refers to the evidence's basic ability to explain the target anomalous event, which can be quantitatively represented by anomaly detection scores, rule matching scores, semantic anomaly confidence, propagation importance scores, and event importance scores. Data source quality can be evaluated by combining sampling completeness, time synchronization degree, field missing rate, data credibility, or collection reliability. Duplicate evidence can be identified and merged or removed based on evidence source, timestamp, evidence type, entity identifier, and summary characteristics. After the above screening of evidence that does not meet the requirements for explanatory power or data source quality, as well as duplicate evidence, further pruning of the context-related evidence set is completed. The resulting set can serve as a candidate evidence set. Figure 2 The flowchart illustrating an example method for determining a set of candidate evidence provided in this application embodiment can be combined with... Figure 2 Understand the above description of the process for determining the candidate evidence set.
[0033] The pre-cutting method will be explained in detail below.
[0034] Let the pre-clipping decision function be:
[0035] when When it indicates that evidence e is pre-cropped; when This indicates that the evidence was discarded during the candidate stage. The pre-pruning decision function can be composed of a combination of factors, such as whether the basic explanatory power of the evidence is below a global or event-level minimum threshold, for example... Whether the data source of the evidence is marked as unstable or abnormal, such as monitoring probe failure or incomplete log collection; whether the evidence is duplicated, such as when multiple semantically equivalent records appear in the same time window and on the same entity, only representative evidence is retained.
[0036] In one implementation, the pre-pruning decision function can also be formalized as:
[0037] in, As an indicator function; For evidence quality assessment functions, such as those calculated based on indicators like sampling integrity, acquisition delay, and parsing error rate; Each of these is a threshold for the basic explanatory power of evidence and a threshold for the quality assessment of evidence in the pre-pruning decision function.
[0038] Using the pre-pruning decision function described above, the technical solution of this application can determine the context-dependent evidence set. The final set of candidate evidence was obtained through screening. :
[0039] In another optional implementation, to control the size of the candidate evidence set and prevent it from far exceeding preset constraints (such as event-level budget constraints) and thus causing excessive subsequent computational resource overhead, a simple budget pre-alignment process can be performed during the candidate evidence set generation stage in step S101. For example, when the size of the candidate evidence set... Far greater than the event-level quantity budget When the upper limit is reached, the evidence can be sorted by basic explanatory power or importance, and only the top K candidate pieces of evidence can be retained in the candidate evidence set, in the following form:
[0040] in this case, This will serve as the set of candidate evidence input for the evidence screening stage in subsequent steps, while the previously obtained evidence... This can be reserved as an alternative set for fallback or expansion if necessary. In optional implementations, this application can directly set... As Whether or not to perform budget pre-alignment processing in this step can be flexibly determined based on the actual resource configuration scale of the operation and maintenance analysis engine, root cause localization system, evidence screening device or processor.
[0041] Through the candidate evidence generation and pre-pruning process described above, this application effectively controls the size of the candidate evidence set while covering important evidence related to the alarm event.
[0042] In summary, during the candidate evidence set generation stage in step S101, a multi-layered screening process combining time constraints, topological constraints, and quality constraints is employed to improve evidence relevance and control the scale of subsequent screening. First, an observation time window is determined based on the alarm time of the target anomaly. Evidence corresponding to data whose time range intersects with the observation time window is extracted from multi-source data, forming a time-based initial screening evidence set. Second, combining the alarm entity set related to the target anomaly and the service and resource topology graph, an alarm context node set is determined. Evidence from the time-based initial screening evidence set whose source entities are located within the alarm context node set is selected, forming a context-related evidence set. Finally, evidence with insufficient explanatory power, unsatisfactory data source quality, and duplicate evidence is removed from the context-related evidence set to obtain the candidate evidence set. Through multi-layered screening, the anomaly relevance is ensured while reducing the introduction of obviously irrelevant evidence.
[0043] Step S102: Based on the explanatory power of the candidate evidence in the candidate evidence set for the target abnormal event, the degree of conflict between the candidate evidence, and the stable support of the candidate evidence set for the root cause conclusion, the evidence is screened from the candidate evidence set to obtain the initial evidence set.
[0044] In step S102, based on the candidate evidence set obtained in the previous step S101, an initial stage of evidence screening is performed, taking into account the explanatory power of the candidate evidence for the target anomalous event, the degree of conflict between candidate evidence in the set, and the stable support of the evidence set for the root cause conclusion. The aim is to obtain a high-quality initial evidence set. In short, in this step, the evidence combination that can better explain the anomalous phenomenon, has fewer conflicts with other evidence, and can provide stable support for the root cause conclusion needs to be selected from the candidate evidence set as the initial evidence set.
[0045] In one example implementation, an incremental iterative screening approach can be used to determine the initial evidence set. Specifically, first, an empty current evidence set is established, and the candidate evidence set is used as the evidence space to be screened. Then, for each candidate piece of evidence in the candidate evidence set, the incremental benefit of adding the corresponding candidate piece of evidence to the current evidence set is evaluated to determine the new evidence set formed.
[0046] Here, incremental gains are used to characterize the impact of adding candidate evidence on the overall explanatory power of the current evidence set, the degree of conflict among evidence within the set, and the set's ability to stably support the root cause conclusion. These impacts may be positive or negative; therefore, in the subsequent screening process, it is necessary to combine the specific numerical values of incremental gains to evaluate which piece of evidence should actually be added to the current evidence set.
[0047] It should be noted that when evaluating incremental gains, it is assumed that a candidate piece of evidence has been added to the current evidence set, rather than that it has actually been added to the set.
[0048] By using incremental gains rather than static scoring of single evidence to measure the value of evidence, the selection logic for candidate evidence can focus on whether adding a piece of evidence to the current set helps to form a better combination of evidence, thus better meeting the actual needs for collaborative interpretation of evidence in root cause localization scenarios.
[0049] When quantifying incremental gains, the overall quality of the current evidence set can be described using an objective function. This objective function can include a first, second, and third calculation term. The first calculation term represents the explanatory power of the evidence set for the target anomalous event; the second calculation term represents the degree of conflict among candidate evidence within the evidence set; and the third calculation term represents the stable support capacity of the evidence set for the root cause conclusion. In the objective function, the first, second, and third calculation terms can be directly related through mathematical operations. Alternatively, they can be processed using weighting coefficients (such as reward coefficients and penalty coefficients) before mathematical operations.
[0050] For any candidate piece of evidence, the difference in the objective function value before and after its addition to the current evidence set can be taken as the corresponding incremental gain. This allows us to measure, on the one hand, the contribution of the evidence to the overall quality of anomaly explanation, and on the other hand, to simultaneously characterize whether it introduces additional conflicts and whether it strengthens the set's stable support for the root cause conclusion. The following is an example of an expression for the objective function:
[0051] Where S represents the set of evidence for which the objective function value needs to be calculated, and e refers to a candidate piece of evidence within the set of evidence S.
[0052] In this formula, the first calculation term This represents the total explanatory power of set S, reflecting the set's ability to explain the target anomalous event; the second calculation term... The third term, R(S), is used to measure the degree of inconsistency of evidence within a set in terms of time, topology, and statistical properties, i.e., the degree of conflict. It is a stability index of the set, reflecting the consistency of the root cause conclusion in multiple replays, thus representing the stable support for the root cause conclusion. The conflict penalty weight is used to adjust the influence of the second calculation term in the objective function; The stability reward weight is used to adjust the influence of the third calculation term in the objective function. The objective function in the above equation can be understood as follows: when selecting evidence, we hope to maximize the explanatory power of the set, while suppressing internal conflicts and encouraging sets that perform more stably in replay.
[0053] Add a candidate piece of evidence The incremental benefit brought about by adding evidence to the evidence set S can be expressed as:
[0054] In some alternative implementations, to reduce computational complexity, the calculation of incremental gains can be approximated, considering only changes in overall interpretability and local conflict. For example, the incremental gains can be approximated as follows:
[0055] in For the comprehensive explanatory power of evidence e, Let be the overall conflict degree of evidence e. The methods for obtaining the overall explanatory power and overall conflict degree will be discussed in detail later and will not be elaborated here. Since the overall explanatory power of evidence is related to its basic explanatory power and its stability factor, the approximate representation of incremental gains in the above formula still considers the impact on the explanatory power of the evidence set, the impact on the degree of conflict between pieces of evidence in the evidence set, and the impact on the stable support capacity of the evidence set for the root cause conclusion.
[0056] To clearly explain the various computational terms in the objective function, a target set is used as an example to introduce the specific computational elements in the first computational term of the target set when it is substituted into the objective function. It should be noted that the target set is used only as a general term in the following description; in specific implementations, the target set substituted into the objective function can be any set of evidence within the technical solution of this application that requires function value quantification through the objective function.
[0057] Baseline explanatory power characterizes the direct contribution of a single piece of evidence to explaining observed anomalies in the current root cause localization task. The higher the value, the more important and explanatory the evidence is considered within the existing root cause localization model or rule system. In some implementations, the basic explanatory power for each candidate piece of evidence can be determined first, based on its evidence type and corresponding explanatory power metric. Evidence types include: indicator-based evidence, log-based evidence, link and topology-based evidence, and change event-based evidence. For example, indicator-based evidence might correspond to anomaly detection scores, log-based evidence to rule matching scores, log template anomaly levels, or semantic anomaly confidence, link and topology-based evidence to propagation importance scores, node anomaly scores, or graph model output scores, and change event-based evidence to event importance scores. For ease of understanding, the basic explanatory power for each type of evidence is given below. The method of determination.
[0058] For indicator-based evidence, such as time series data like service latency, error rate, and CPU utilization, this application can utilize the anomaly detection score output by the anomaly detection module as the basic explanatory power. Let the anomaly score corresponding to a certain indicator-based evidence be... Then you can directly order In one implementation method, to facilitate comparison of different types of evidence, the anomaly detection score can be normalized according to historical statistics or distribution characteristics, making... It falls within a preset range, such as [0, 1].
[0059] For log-based evidence, this application can use the rule matching results, the anomaly level of the log template, or the anomaly confidence score output by the log semantic analysis model as the basic explanatory power. For example, suppose the log analysis model gives an anomaly confidence score of 1 / e for a certain log record. Then it can be defined When logs are triggered by simple rules, fixed weights or score mapping based on rule priority can also be used. .
[0060] For link and topology-related evidence, such as trace spans, service call edges, or resource dependency edges, this application can utilize the node or edge anomaly scores output by the topology propagation model as the basic explanatory power. Let the propagation importance score corresponding to a certain link evidence e be . Then there is In other implementations, the PageRank value of the service node, the failure propagation probability, or the node anomaly score output by the graph neural network can be mapped to the basic explanatory power of a single link or node evidence; this application does not limit this approach.
[0061] For evidence of change events such as releases, configuration modifications, scaling up or down, this application can construct a change importance score based on factors such as the type of change, the proximity of the change time to the alarm time, and the topological distance between the changed object and the abnormal service. and order .
[0062] In optional implementations, multiple influencing factors can be integrated into a single scalar score through simple linear or nonlinear combinations.
[0063] The optional implementation method for calculating the basic explanatory power is as follows: since the basic scores of different types of evidence may come from different models and dimensions, this application, after completing the calculation of the basic explanatory power of each type of indicator, can apply it to all... Perform normalization or stretching transformations. For example, the underlying explanatory power of each piece of candidate evidence in the candidate evidence set can be linearly normalized:
[0064] in, and Let represent the minimum and maximum values of the basic explanatory power in the current candidate evidence set, respectively. To prevent extremely small positive numbers with a denominator of zero, normalization is used. It can replace the originally calculated basic explanatory power. This is used for the calculation of subsequent explanatory power-related indicators and evidence selection, thereby ensuring that the explanatory power of different data sources and model outputs is within a comparable numerical range.
[0065] In summary, the basic explanatory power in this application is... This is a unified abstraction of the importance signals output by existing root cause localization models or rule engines. To explicitly quantify whether evidence makes the root cause conclusion more stable, this application introduces a stability index for the evidence set and a stability factor for individual pieces of evidence, based on multiple replays, and incorporates these as components of the subsequent calculation of the overall explanatory power. Using the stability factor and the previously introduced basic explanatory power, the overall explanatory power can be calculated. The method for obtaining the stability factor of candidate evidence will be described next.
[0066] Specifically, for each candidate piece of evidence in the target set, a pruned subset corresponding to that candidate piece of evidence is constructed, and stability indices are obtained for both the pruned subset and the target set. The pruned subset, compared to the target set, only lacks the corresponding candidate evidence.
[0067] This application defines the root cause conclusion obtained from the i-th playback when the same set of evidence S is played N times independently as:
[0068] In one implementation method, the result of the first playback can be selected. As a baseline conclusion, a stability index for set S is defined. for:
[0069] in Let be an indicator function. This function takes a value of 1 when the i-th replay is completely consistent with the baseline conclusion, and 0 otherwise. Through this definition, the replay consistency of the evidence set S can be quantified to the interval [0,1]. The larger the value, the more stable the root cause conclusion is in multiple replays.
[0070] In another implementation, to allow for slight differences within a certain range, this application may also introduce a distance metric function between root cause conclusions. If two conclusions measured by this function are completely consistent, the distance is considered to be 0. If the difference between the two conclusions is positively correlated with the distance, then the greater the difference, the greater the distance is assigned. In this case, the stability index of set S can be... Defined as:
[0071] in, The value of is normalized to the interval [0,1], for example, based on the Jaccard distance of the root cause node set or the rank distance based on the root cause sorting list. This application may choose any of the above forms to calculate the stability index of the set.
[0072] By using any of the above methods, the stability index corresponding to the abridged subset and the stability index corresponding to the target set can be calculated. Applying the same calculation method, only the calculated sets show slight differences in evidence.
[0073] To characterize the contribution of a single piece of evidence to the stability of the root cause conclusion, this application introduces a single-evidence stability factor. Here, e represents a piece of evidence. Calculating the stability factor, simply put, involves comparing "using a set of evidence S" with "the set of S after removing evidence e". "The difference in stability between the two cases. The former can be understood as the target set, while the latter can be understood as the abridged subset of evidence e."
[0074] In one alternative implementation, the same set of evidence S and the abridged subset corresponding to evidence e are considered. After performing N replays, two sets of stability indices were obtained, as shown below:
[0075] Then the stability factor of evidence e is defined as:
[0076] This definition reflects the degree to which the overall replay stability of the evidence set S decreases after removing evidence e. When removing a piece of evidence e significantly reduces the stability of the set, At this point, stab(e) takes a positive value, indicating that evidence e makes a positive contribution to maintaining the stability of the root cause conclusion; however, when removing evidence e has little impact on stability or even improves stability, based on the above formula, the stability factor stab(e) is truncated to 0, which means that the evidence is not important from the perspective of stability or even has a negative impact.
[0077] In another implementation, the frequency of "whether removing 'e' changes the root cause conclusion" can be directly counted from the playback results. For example, for each playback i, the following can be recorded:
[0078] Define the influence indicator variable:
[0079] As an indicator function, if the values within the function are the same, the function value is 1; otherwise, it is 0. Therefore, another form of stability factor can be defined:
[0080] In this form, the larger the stability factor stab(e), the more times the evidence e is removed will cause the root cause conclusion to change, indicating that evidence e has a higher structural influence on the conclusion.
[0081] In the technical solution of this application, either of the two aforementioned calculation methods or a combination of the two can be used to calculate the stability factor.
[0082] As introduced above regarding the first technical implementation of the stability factor, the stability factor of the candidate evidence can be determined based on the difference between the stability index of the pruned subset and the corresponding stability index of the target set, thus reflecting the degree of contribution of the candidate evidence to the stability of the set.
[0083] Subsequently, combining the basic explanatory power and the stability factor, the overall explanatory power of each candidate piece of evidence is determined, and the sum of the overall explanatory power of the target set is obtained by aggregating them, serving as the first calculation item. This approach allows the evaluation of explanatory power to move beyond the static quality of individual pieces of evidence and simultaneously incorporate the actual contribution of candidate evidence to the stability of the set. To simultaneously consider the basic importance given by existing root cause localization models and the contribution of evidence to the stability of the conclusion, this application uses the overall explanatory power of a single piece of evidence e... Defined as:
[0084] in, The basic explanatory power, as introduced earlier, can be provided by existing models or rule engines; The stability factor for a single piece of evidence, used to characterize the structural impact of that evidence on the stability of the replay of the entire collection, has already been introduced earlier; These are adjustable weights used to control the relative importance of the stability factor in the overall explanatory power.
[0085] In one type of implementation, it is possible to... and After normalizing each term and constructing the comprehensive explanatory power in a linear combination form, the above formula can be considered a simplified representation of this combination form. This application does not limit the specific normalization and combination methods, as long as the comprehensive explanatory power g(e) can comprehensively reflect both "explanatory power" and "stability contribution." Based on the preceding introduction and description of the objective function, it is not difficult to find that the first calculation term in the objective function... It is essentially a summation, adding up the comprehensive explanatory power of each piece of evidence in the target set (e.g., the evidence set S).
[0086] Next, for ease of explanation, we will take the target set as an example to introduce the specific calculation elements of the second calculation term of the target set when it is substituted into the objective function.
[0087] For the second calculation item, in some implementations, the conflict of evidence can be characterized from three dimensions: time, topology, and statistics. Specifically, the degree of temporal conflict reflects the degree of deviation of a single candidate piece of evidence from the main timeline of the set or the time representing the set; the degree of topological conflict reflects the degree of deviation of the source entity of the candidate evidence from the main propagation subgraph, the path deviation from the candidate root cause node, or the inconsistency of the propagation direction; and the degree of statistical conflict reflects the degree of inconsistency between the direction, magnitude, or abnormal pattern of the indicator change corresponding to the candidate evidence and the mainstream evidence in the same group. Based on at least two of the above three conflict dimensions for each candidate piece of evidence, the comprehensive conflict degree of the evidence can be calculated, and then the total conflict degree of the target set can be obtained as the second calculation item.
[0088] Understandably, in multi-source evidence scenarios, even if a single piece of evidence has high basic explanatory power, if it is significantly inconsistent with other evidence in terms of chronological order, topological path, or statistical characteristics, then that evidence is likely to be noise or a coincidental association. Therefore, this application uses a multi-source conflict measurement method to characterize the degree of inconsistency between a single piece of evidence and a current set of evidence.
[0089] In one example, for the set of evidence One piece of evidence This application defines component conflict degree from three aspects: time, topology, and statistics, and combines them into a comprehensive conflict degree c(e,S). The calculation methods for time conflict degree, topology conflict degree, and statistical conflict degree are introduced below.
[0090] To measure the degree of inconsistency of evidence over time, this application first defines a representative time point for the evidence, for example, taking the midpoint of its time interval:
[0091] For a set of evidence S, a representative time statistic can be defined for the set, such as the median or a weighted average time point. Among the optional implementations, the median form is used:
[0092] The degree of deviation of a single piece of evidence e in the time dimension can be defined as:
[0093] If a time-consistent budget exists, then in order to align with the time-consistent budget... By aligning the time discrepancies, this application can normalize the time discrepancies to obtain the degree of time conflict:
[0094] in To prevent extremely small positive numbers with a denominator of zero, when the time of a piece of evidence deviates significantly from the main time period of the set, A value close to 1 indicates a strong conflict in the time dimension.
[0095] In cloud-native systems, there are calling and dependency relationships between services and resources. This application uses a directed graph. Represents the service and resource topology, where For a set of nodes, Let be the set of edges. The source entity src(e) of evidence e can be mapped to a node or an edge on the graph.
[0096] Let the set of candidate root cause nodes in this root cause localization task be . Based on the overall distribution of the evidence set S, the main propagation path or main propagation subgraph can be deduced. This subgraph reflects the anomalous propagation direction supported by most of the evidence. For a single piece of evidence e, this application characterizes its degree of topological conflict by considering whether the source entity src(e) is related to... The nodes in the array are connected; src(e) is connected to the set of candidate root cause nodes. Is the path length significantly excessive? Is the propagation direction supported by evidence (e.g., from upstream to downstream, or from resources to services) opposite to the main propagation direction?
[0097] In one implementation approach, the degree of topological conflict can be defined as:
[0098] in, Main propagation subgraph The set of nodes; This is an indicator function that generates a fundamental conflict when the evidence source node is not in the main propagation subgraph; It is a graph distance-based normalization function used to characterize the deviation of the shortest path length from src(e) to the main propagation path or the set of candidate root cause nodes; These are adjustable weights used to balance the contributions of the two parts.
[0099] In other implementations, the degree of topological conflict can also be constructed using indicators such as edge attention weights and structural similarity output by the graph neural network, and this application does not limit this.
[0100] Statistical conflict primarily addresses the characteristics of metrics and log values, characterizing contradictions between pieces of evidence from the perspective of consistency in the direction and magnitude of data changes. For example, if the vast majority of metrics related to a certain service indicate a significant increase in latency, while a particular piece of evidence shows a decrease or no change in latency, then this piece of evidence statistically conflicts with the mainstream evidence. In optional implementations, this application first clusters or groups the evidence related to the same metric or measure in the evidence set S, obtaining a grouped set. For the group containing single piece of evidence e The degree of statistical conflict can be defined as:
[0101] in, The direction of change (e.g., rising, falling, or remaining unchanged) of the indicator or metric corresponding to evidence e can be represented by discrete values. It indicates the direction of change in the dominant evidence within a group, for example, through majority voting or weighted statistics; This is a direction conflict function; its value is close to 0 when the two directions are the same, and close to 1 when they are opposite. In a simple implementation, it can take the following values:
[0102] In more complex implementations, the magnitude of change and the degree of anomaly can be considered together, so that the greater the deviation of the direction and magnitude of change from the mainstream, the greater the degree of statistical conflict.
[0103] To comprehensively consider conflicts across the three dimensions of time, topology, and statistics within a unified framework, this application combines these conflict levels in a weighted manner, defining the comprehensive conflict degree of a single piece of evidence e as:
[0104] in These are adjustable weights, used to adjust the importance of each dimension of conflict in the overall conflict level according to business needs.
[0105] The total conflict degree of the evidence set S is defined in this application as the sum of the combined conflict degrees of all evidence within the set:
[0106] Conflict budget can be used The upper bound constraint given by the total conflict degree C(S) of the set:
[0107] By using the aforementioned multi-source conflict measurement method, this application can identify and suppress evidence that is significantly inconsistent with the main time period, main propagation path, or main statistical characteristics during the selection of evidence subsets, thereby avoiding individual high-scoring but highly conflicting evidence from misleading the root cause conclusions and providing a foundation for constructing a minimal and stable chain of evidence under preset constraints.
[0108] The third computational term, namely the stability index of the set, has already been introduced in the introduction of the first computational term, so it will not be repeated here.
[0109] In practical applications, a stability threshold (lower stability constraint) can be preset or provided by the policy learning module to exclude evidence sets that are unstable in overall playback. In other words, the stability index of the set should be greater than or equal to the stability threshold. Otherwise, the playback stability of the set can be determined to be insufficient.
[0110] Furthermore, in some example implementations, the filtering process can be performed under preset constraints, thereby better controlling the size of the set and computational complexity.
[0111] The preset constraints can be event-level budget constraints, used to limit the budget indicators allowed to be included in a single root cause localization task. Based on this, the incremental benefits are calculated only for candidate evidence addition schemes that meet the preset constraints, and the target candidate evidence is determined to be actually added to the current evidence set. In specific implementation, the candidate evidence with the largest incremental benefit or multiple candidate evidences with large incremental benefit values can be added to the current evidence set.
[0112] After each addition operation, the incremental gain of each remaining candidate piece of evidence needs to be recalculated based on the updated current evidence set and the remaining candidate evidence. This is because the value of a candidate piece of evidence is not fixed but depends on the existing evidence in the current evidence set. For example, a piece of evidence might provide a high explanatory gain in an empty set, but its marginal gain might decrease after similar highly relevant evidence is added. Similarly, if the current evidence set already contains key evidence on a certain topological path, adding conflicting evidence could significantly increase the overall conflict level. Therefore, re-evaluating the remaining candidate evidence after each update of the current evidence set helps to dynamically approximate the globally optimal evidence combination. It should be noted that after candidate evidence is added to the current evidence set, the corresponding candidate evidence can be removed from the candidate evidence set; alternatively, the status of each candidate piece of evidence in the candidate evidence set can be used to distinguish whether it has been selected for addition to the current evidence set.
[0113] For constructing the initial evidence set in this application, a greedy strategy or other approximate strategies such as bundle search can be used. In one implementation, this application uses a greedy strategy to gradually construct the initial evidence set under budget constraints. The algorithm initially sets the current evidence set as follows:
[0114] Then repeat the following steps until no more new evidence can be added within the budget constraint: From the remaining candidate evidence in the candidate evidence set, select the set of evidence that, when added to the current evidence set, still ensures that the updated current evidence set satisfies the preset constraints. ;like If the result is empty (i.e., no such evidence exists), then the iteration terminates; otherwise, it means there is still evidence that satisfies the requirements in 1) above, then for each Calculate approximate incremental returns ( The evidence with the largest incremental return is selected as follows:
[0115] Next Add to the current evidence set, in the following form:
[0116] In a greedy algorithm, each iteration selects the evidence that maximizes the incremental gain under preset constraints, until the addition of any candidate evidence would violate the preset constraints or the incremental gain would no longer be positive. The final result is... A feasible set of evidence that has a high objective function value under preset constraints can be used as the initial set of evidence.
[0117] In another implementation, considering that a single-path greedy strategy may get stuck in local optima in certain scenarios, this application can also employ multi-path expansion strategies such as beam search to improve solution quality within acceptable computational overhead. Taking beam search as an example, assuming the beam width is M, this application maintains a current evidence set list:
[0118] In each iteration, the bundle search performs a search on each set in the list. Extend one or more pieces of evidence according to the above definition of incremental revenue to generate several new sets. Then, select the top M sets from all new sets as the bundle list for the next round based on the objective function F(S) or its approximation. The stopping condition for the bundle search can include budget exhaustion, insufficient improvement of the objective function, or reaching the upper limit of the number of iterations. Finally, select the set with the largest objective function value from the bundle list as the bundle list. .
[0119] In some implementations, when there are no remaining candidate pieces of evidence in the candidate evidence set that satisfy the preset constraints, or when the incremental gains corresponding to the remaining candidate pieces of evidence do not satisfy the preset constraints, the current iteration can be stopped, and the current evidence set obtained from the last update can be used as the initial evidence set. Here, the preset constraints refer to numerical constraints on the incremental gains. For example, the preset constraint is that the incremental gain is greater than a value K, where K is used to define whether the incremental gain is too low. If the preset constraint is not met, it is considered that the incremental gain of a certain piece of evidence is too low and is not suitable for selection into the current evidence set.
[0120] The initial evidence set obtained in step S102 has higher quality of anomaly interpretation and better evidence synergy compared to the candidate evidence set, thus contributing to improving the stability and efficiency of the overall root cause localization process. Although it has been optimized in terms of interpretation quality, conflict control, and stable support capabilities, it may still contain redundant evidence that is not necessary for the root cause conclusion. Therefore, further counterfactual pruning is required in subsequent steps S103-S105 to further refine the evidence set.
[0121] Step S103: For each candidate piece of evidence in the initial evidence set, construct a counterfactual evidence set after deleting the corresponding candidate piece of evidence, and determine whether the counterfactual evidence set meets the preset stability requirements and preset explanatory power requirements by replaying the evidence.
[0122] In step S103, counterfactual analysis is further performed on the initial evidence set obtained in the previous step S102 to determine whether each candidate piece of evidence is necessary to maintain the stability of the root cause conclusion. By attempting to delete candidate evidence one by one and observing the performance of the deleted evidence set in the playback scenario, redundant and unnecessary evidence in the initial evidence set is identified, thereby providing a basis for obtaining a more compact and representative target evidence set.
[0123] Specifically, each candidate piece of evidence in the initial evidence set can be iterated through one by one. For any candidate piece of evidence to be judged in the initial evidence set, the candidate piece of evidence is deleted from the initial evidence set, and a corresponding counterfactual evidence set is constructed.
[0124] A counterfactual evidence set refers to a set of evidence that, relative to the initial evidence set, is missing only one specific candidate piece of evidence to be determined. By constructing this counterfactual evidence set, we can simulate the scenario of "whether the current root cause localization process can still obtain a stable and sufficiently credible result if this candidate piece of evidence is not used," thereby verifying the necessity of the candidate piece of evidence.
[0125] In terms of specific implementation, the counterfactual evidence set can be replayed. Replay can be understood as taking the counterfactual evidence set as input and re-executing the root cause localization analysis process corresponding to the target anomaly, or recalculating and verifying the root cause localization conclusions in preset historical scenarios, simulation scenarios, or offline diagnostic scenarios. Through replay, the stability index and objective function value corresponding to the counterfactual evidence set can be obtained.
[0126] Stability metrics have been introduced previously. Stability metrics measure the consistency of the root cause conclusions corresponding to a set of counterfactual evidence under multiple replays, different sampling perturbations, different sub-scenes, or different random initialization conditions. In some implementations, stability metrics can be determined by the consistency of root cause conclusions obtained after multiple independent replays. For example, it can be calculated by statistically analyzing whether the entities identified as root cause nodes are consistent across multiple replays, whether the order is stable, and whether the core propagation path is consistent, and then calculating a consistency score accordingly.
[0127] The objective function, as previously described, comprehensively characterizes the counterfactual evidence set's explanatory power for the target anomaly, the degree of conflict among the evidence, and its stable support for the root cause conclusion. The objective function value can follow the framework of the objective function mentioned in step S102, determined through a combination of the first, second, and third calculation terms. Specifically, the first calculation term represents the counterfactual evidence set's explanatory power for the target anomaly; the second calculation term represents the degree of conflict among candidate evidence within the counterfactual evidence set; and the third calculation term represents the counterfactual evidence set's stable support for the root cause conclusion.
[0128] When the stability index corresponding to the counterfactual evidence set is not lower than a preset stability threshold, and the objective function value is not lower than a preset explanatory power threshold, it can be determined that after deleting the candidate evidence to be judged, the counterfactual evidence set corresponding to that evidence can still meet the requirements in terms of explanatory quality and conclusion stability. In other words, the candidate evidence is not necessary to maintain the stable validity of the current root cause conclusion. Conversely, if the stability index drops below the preset stability threshold after deletion, and / or the objective function value is lower than the preset explanatory power threshold, it indicates that the candidate evidence plays an irreplaceable role in maintaining the overall quality of the set and should be retained, without being deleted from the initial evidence set.
[0129] In this embodiment, step S103 uses a counterfactual approach to examine the necessity of a single candidate piece of evidence in the evidence set. It is important to note that this process not only focuses on whether the result is still acceptable, but also on whether the evidence set after deletion remains within preset ranges in terms of both stability and interpretive quality. This provides a basis for subsequent actual deletion operations and makes the resulting target evidence set more likely to be the core evidence set that stably supports the root cause conclusion.
[0130] Step S104: If the counterfactual evidence set meets the preset stability requirements and preset explanatory power requirements, delete the corresponding candidate evidence from the initial evidence set.
[0131] In step S104, based on the replay judgment result of step S103, candidate evidence in the initial evidence set is deleted. This step transforms the counterfactual judgment conclusion into an actual evidence set update operation, that is, removing candidate evidence that is not necessary for the stable establishment of the root cause conclusion from the current evidence set, thereby reducing the size of the initial evidence set and improving the core and interpretability of the initial evidence set.
[0132] Specifically, when the counterfactual evidence set formed after the deletion of a candidate piece of evidence simultaneously meets the preset stability and explanatory power requirements, it indicates that even in the absence of that candidate piece of evidence, the current evidence set can still provide sufficient explanation for the anomaly and still provide stable support for the root cause conclusion. In this case, the candidate piece of evidence can be deleted from the initial evidence set, and the deleted evidence set can be used as the new set to continue the deletion judgment of subsequent candidate pieces of evidence. Through this process, unnecessary redundant evidence can be gradually eliminated, reducing evidence items with repeated expressions, weak correlations, or low marginal contributions in the evidence set.
[0133] In some implementations, the deletion strategy can employ either sequential deletion or deletion based on a preset order. For example, candidate evidence in the initial evidence set can be sorted according to factors such as the basic explanatory power, contribution to the objective function, stability factor, conflict contribution, or chronological order, and then counterfactual determination and deletion can be performed sequentially. Regardless of the order, the core principle is that candidate evidence is only truly deleted if the counterfactual evidence set formed after deleting a candidate evidence still meets the preset stability and explanatory power requirements. This ensures that the deletion process does not compromise the overall validity of the set.
[0134] Furthermore, during the deletion operation, various statistics or evaluation indicators related to the initial evidence set can be updated simultaneously. For example, the conflict relationship diagram between candidate evidence within the initial evidence set, the total conflict degree of the set, the comprehensive explanatory power of the set, and the cached or replay records of stability indicators can be updated to facilitate subsequent candidate evidence deletion decisions. By dynamically updating the evaluation basis, subsequent judgments can more accurately reflect the actual impact of the deletion operation on the evidence set.
[0135] In this embodiment of the application, step S104 can remove the unnecessary candidate evidence identified in step S103 from the initial evidence set, so that the evidence set can continuously converge towards a more compact, more stable and more explanatory direction.
[0136] Step S105: Repeat the candidate evidence deletion judgment until there is no candidate evidence that still meets the preset stability requirements and preset explanatory power requirements after deletion, and obtain the target evidence set.
[0137] In step S105, based on the updated evidence set after the deletion operation in step S104, the candidate evidence deletion judgment is repeated until there are no candidate evidences in the initial evidence set that still meet the preset stability and preset explanatory power requirements after deletion, thus obtaining the final target evidence set. It can be understood that this step extends the single counterfactual judgment and deletion mechanism formed in steps S103 and S104 into an iterative convergence process, ensuring that every piece of retained evidence in the target evidence set is necessary or important to the overall quality of the set.
[0138] In practice, after each deletion operation, the updated initial evidence set can be used as input again to iterate through the candidate evidence, constructing a new counterfactual evidence set. The stability index and objective function value after deleting the corresponding candidate evidence can then be re-evaluated through replay. Since the marginal effect of the remaining candidate evidence in the set may change after deleting a candidate evidence, the judgment needs to be re-executed on the updated initial evidence set. Using an iterative deletion strategy can avoid the bias caused by a one-time static judgment and improve the rationality of the final target evidence set.
[0139] In some implementations, the stopping condition for step S105 may include at least one of the following: First, after any candidate evidence in the current evidence set is deleted, the resulting counterfactual evidence set cannot simultaneously meet the preset stability requirement and the preset explanatory power requirement; second, the current initial evidence set has been reduced to a preset minimum size through pruning; third, the set no longer changes in several consecutive iterations. When the stopping condition is met, the current initial evidence set that has undergone multiple prunings can be determined as the target evidence set.
[0140] Compared to the initial evidence set obtained in step S102, this target evidence set further eliminates candidate evidence that is not necessary for the stable supporting root cause conclusion, and therefore usually has higher explanatory compactness and stronger conclusion stability.
[0141] Existing technologies typically generate lengthy evidence sets or chains, lacking stability verification mechanisms. Due to the large volume, complex sources, and frequent fluctuations of data in cloud-native systems, the root cause localization conclusions may change under different times, sampling conditions, or execution batches for the same anomaly. This results in poor reproducibility and insufficient credibility of diagnostic results, hindering operational review and audit trails. In multi-source evidence fusion scenarios, existing solutions often focus on information overlay without explicit quantification and constraints on evidence conflicts. For example, some evidence may deviate from the main anomaly period in time, the main propagation path in topology, or the direction of statistical characteristics opposite to most anomalies, yet it may still be included in the final result due to its high local anomaly score, thus misleading root cause analysis. Furthermore, existing technologies generally lack counterfactual verification mechanisms for the necessity of individual pieces of evidence, making it difficult to determine whether a piece of evidence is a necessary condition for the stable establishment of a root cause conclusion. Consequently, it is impossible to effectively eliminate weak evidence that is "dispensable," making it difficult to output a concise and stable evidence set, and even more difficult to construct a clearly structured evidence chain that is easy to present and review.
[0142] In this embodiment, evidence screening is not based on the static anomaly score of a single piece of evidence. Instead, after constructing a candidate evidence set for the target anomaly event, a set-level screening is performed based on the synergistic relationships between the evidence. Specifically, on the one hand, the explanatory power of the candidate evidence for the target anomaly event is evaluated, with explanatory power characterizing the degree to which each candidate piece of evidence explains the anomaly; on the other hand, the degree of conflict between the candidate evidence is evaluated, with the degree of conflict characterizing the inconsistency between the evidence in terms of time, topological location, statistical characteristics, or propagation direction; simultaneously, the stable support capability of the candidate evidence set for the root cause conclusion is further evaluated, with stable support capability measuring the degree of consistent support of the evidence set for the root cause conclusion during playback or repeated diagnosis. Through the above multi-dimensional joint evaluation, an initial evidence set can be obtained from the candidate evidence set. Subsequently, by constructing a counterfactual evidence set and combining it with playback verification, it is determined whether the candidate evidence in the set is necessary to maintain the stable validity of the current root cause conclusion, thereby deleting unnecessary evidence and obtaining the target evidence set. This forms a two-stage evidence screening mechanism consisting of initial screening and counterfactual pruning. This effectively reduces the inclusion of redundant, noisy, and conflicting evidence, thereby improving the accuracy, interpretability, and reproducibility of root cause localization results by leveraging the target evidence set.
[0143] In some implementations, after obtaining the target set of evidence, further evidence chain construction processing can be performed. Figure 3 This is a flowchart illustrating the construction of an evidence chain based on a target set of evidence, provided as an embodiment of this application. The steps shown in the flowchart can be specifically described in... Figure 1 This is executed after step S105. (Combined with...) Figure 3 The evidence chain construction process includes: determining the abnormal related nodes covered by the evidence based on the target evidence set, and extracting the corresponding connected subgraph from the service and resource topology graph; determining the propagation order of the abnormal related nodes in the connected subgraph according to the time information corresponding to each piece of evidence in the target evidence set; searching for the propagation path from the root cause candidate node to the affected node in the connected subgraph, and selecting representative evidence corresponding to the path node along the propagation path; and constructing the minimum evidence chain based on the propagation path, propagation order, and representative evidence.
[0144] Understandably, this final minimal chain of evidence can be used to help demonstrate the anomaly propagation process and explain root cause localization results. For ease of understanding... Figure 3 The specific execution details of each step are shown below. The process of constructing the minimum chain of evidence will be described in detail below.
[0145] To obtain the minimum stable evidence set, i.e. the target evidence set Subsequently, this application further constructs a time-continuous, topologically connected, and concise evidence chain on the service and resource topology to demonstrate to operations and maintenance personnel the critical propagation path from the root cause to the observed anomaly. Compared to simply listing evidence sets, the evidence chain format can more intuitively reflect the process of a fault propagating from the root cause node through several key nodes to the affected service, facilitating understanding and review.
[0146] First, we introduce how to construct a subgraph based on the target evidence set. Let the service and resource topology of the cloud-native system be:
[0147] Where V is the set of nodes and E is the set of edges. Based on the target evidence set... This application first constructs a subgraph supported by evidence on the topological graph. Let it be... The set of nodes covered is:
[0148] Here, src(e) represents the source entity of evidence e. In one implementation, only the entity derived from the original topology G can be extracted. The resulting induced subgraph:
[0149] in, This subgraph It describes all service and resource nodes supported by the target evidence set and the topological connections between them.
[0150] Next, we introduce the root cause candidate nodes and the set of affected nodes. From the root cause localization results, this application obtains a set of root cause candidate nodes, denoted as:
[0151] This set can be given by the internal inference results of the objective function or other root cause scoring models. Simultaneously, based on the alarm event and business priority, a set of affected nodes can be identified, such as nodes on services directly involved in the alarm and critical business links, denoted as:
[0152] In constructing the chain of evidence, this application prioritizes evidence from... Starting from the root cause candidate node, arrive at The propagation path of the affected nodes.
[0153] Next, the chronological ordering and selection of key nodes are introduced. To ensure the temporal plausibility of the chain of evidence, this application utilizes representative time points of the evidence. Target evidence set Sort the data. For each node... Its representative time can be defined as the summation statistic of the midpoint of all evidence times at that node, such as the median:
[0154] Subsequently, on the set according to The nodes are sorted in non-descending order to obtain a time-ordered sequence. Combined with the topological structure, this application can further filter out "key propagation nodes," such as: root cause nodes and their adjacent upstream or downstream nodes; intermediate nodes on the main propagation path; and nodes supported by multiple pieces of evidence with high comprehensive explanatory power.
[0155] In one implementation, the node-level comprehensive interpretability can be calculated for each node v:
[0156] Combination Based on the node metric G(v), select a set of key nodes. As the main site in the chain of evidence.
[0157] In the root cause node set and key node set Based on this, this application in the subgraph Search for a connected path from the root node to the affected node. For each pair (r, a), where:
[0158] This application may be submitted in [location]. The following calculations are performed to obtain the node sequence: At each path node The above is selected from the most representative or explanatory evidence in terms of time. This serves as representative evidence for that node in the chain of evidence. Therefore, the corresponding chain of evidence can be constructed:
[0159] in, Furthermore, adjacent evidence nodes in the evidence chain are connected by edges in the topological graph. In cases where multiple paths exist, the path with the lowest overall cost can be selected as the final evidence chain based on factors such as path length, the cumulative value of the node's comprehensive explanatory power G(v), and the total conflict degree of the evidence along the path.
[0160] In situations where multiple root causes and multiple affected nodes coexist, this application can construct multiple chains of evidence. And by merging or folding shared prefixes, an evidence tree is formed that extends from the root node to downstream services, and is still presented as a single chain or a small number of links when displayed.
[0161] To further improve the readability of the chain of evidence, this application, after constructing... Subsequently, the following compression and display optimizations can be performed: For multiple pieces of evidence on the same node v that are very close in time and semantically similar, they can be merged into one step in the link, and only a summary of representative evidence can be displayed; when some intermediate nodes on the path only play a forwarding role and the evidence is weak, they can be folded in the link display according to the business configuration, and only expanded when in-depth analysis is required; during the display, each evidence node on the link is arranged in chronological order, and its type (log, metric, link, change event), comprehensive explanatory power, topological distance from the root cause, and other key information are marked to help operations personnel quickly understand the fault propagation process. Finally, the minimum stable evidence chain (or evidence tree) obtained in this application not only ensures topological connectivity and temporal consistency, but also minimizes redundant evidence under the constraints of stability and explanatory power, thereby providing operations personnel with a short, reliable, easily reproducible, and auditable root cause evidence path within a limited display space.
[0162] The evidence screening method and related evidence chain construction scheme proposed in this application for root cause localization establish a joint measurement and phased optimization mechanism around the explanatory power of evidence, the degree of conflict of evidence, and the stable support capability of the conclusion. Specifically, for a single piece of evidence, a comprehensive explanatory power composed of the basic explanatory power and the stability factor obtained based on replay is used for evaluation; for a multi-source evidence set, a conflict measurement function is constructed to characterize the degree of inconsistency of evidence in terms of time, topology, and statistical properties, and a replay stability index is defined at the set level to evaluate the support capability of the evidence set for the stability of the root cause conclusion.
[0163] Based on this, under multi-dimensional budget constraints such as quantity, type, time consistency, and conflict, this application uses an objective function to jointly weigh the comprehensive explanatory power, conflict penalty, and stability reward to generate an initial evidence set. Furthermore, by combining a counterfactual pruning strategy driven by multiple replays, redundant evidence is deleted one by one while meeting the stability and explanatory power requirements, resulting in a minimum stable evidence set (i.e., the target evidence set), and the corresponding minimum stable evidence chain can be further constructed.
[0164] Meanwhile, this application establishes a three-tiered screening process: candidate evidence generation and pre-pruning, subset selection under budget constraints, and counterfactual pruning based on playback. First, coarse pruning is performed along the time and topological dimensions; then, set optimization is completed under budget constraints; and finally, unnecessary evidence is removed through playback. This significantly compresses the evidence size and evidence chain length while maintaining explanatory power and stability. In this way, it avoids the forced correlations and overfitting to noise caused by endlessly piling up evidence, reduces the processing of invalid evidence and unnecessary playback calculations, lowers the computational and storage overhead of online diagnosis, and improves the accuracy, real-time performance, and resource utilization efficiency of root cause localization in large-scale cloud-native platforms and computing power network scenarios.
[0165] Based on the root cause localization-oriented evidence screening method described in the foregoing embodiments, this application also provides a root cause localization-oriented evidence screening device. The technical implementation of this device will be described below with reference to the accompanying drawings. Figure 4 This is a schematic diagram of an evidence screening device for root cause localization provided in an embodiment of this application.
[0166] like Figure 4 As shown in the embodiments of this application, the evidence screening device for root cause localization includes: The evidence set acquisition module 41 is used to acquire multi-source data related to the target abnormal event and determine a candidate evidence set based on the multi-source data; The first screening module 42 is used to screen evidence from the candidate evidence set based on the explanatory power of the candidate evidence in the candidate evidence set for the target abnormal event, the degree of conflict between the candidate evidence, and the stable support power of the candidate evidence set for the root cause conclusion, to obtain an initial evidence set. The counterfactual pruning module 43 is used to construct a counterfactual evidence set after deleting the corresponding candidate evidence for each candidate evidence in the initial evidence set, and to determine whether the counterfactual evidence set meets the preset stability requirements and preset explanatory power requirements by replaying the evidence. If the counterfactual evidence set meets the preset stability requirements and preset explanatory power requirements, the corresponding candidate evidence is deleted from the initial evidence set. The candidate evidence deletion judgment is repeated until there is no candidate evidence that still meets the preset stability requirements and preset explanatory power requirements after deletion, thus obtaining the target evidence set.
[0167] Optionally, the first filtering module 42 specifically includes: An empty set creation unit is used to create an empty current evidence set; The incremental benefit calculation unit is used to calculate the incremental benefit formed by adding the corresponding candidate evidence to the current evidence set for each candidate evidence in the candidate evidence set. The incremental benefit is used to characterize the impact of adding the corresponding candidate evidence. The impact includes: the impact on the explanatory power of the current evidence set, the impact on the degree of conflict between the evidence in the current evidence set, and the impact on the stability support capability of the current evidence set. An evidence addition unit is used to determine target candidate evidence to be added to the current evidence set based on the incremental benefit. The incremental benefit calculation unit is further configured to continue calculating the incremental benefit corresponding to each remaining candidate evidence based on the remaining candidate evidence after the target candidate evidence is added to the current evidence set and the updated current evidence set; the evidence addition unit is further configured to repeatedly perform candidate evidence selection, continuously update the current evidence set, and use the set obtained from the last update as the initial evidence set.
[0168] Optionally, the incremental revenue calculation unit is specifically used for: For each candidate piece of evidence in the candidate evidence set, determine whether the updated current evidence set satisfies the preset constraints after adding the corresponding candidate piece of evidence to the current evidence set. For the updated current evidence set that satisfies the preset constraints, calculate the resulting incremental revenue respectively; The incremental revenue calculation unit is specifically used to: calculate the incremental revenue corresponding to each remaining candidate evidence under the preset constraints, based on the remaining candidate evidence after the target candidate evidence is added to the current evidence set and the updated current evidence set; The evidence addition unit is specifically used to repeatedly perform candidate evidence selection and continuously update the current evidence set until there are no remaining candidate evidences in the candidate evidence set that meet the preset constraints, or until the incremental benefits corresponding to the remaining candidate evidences do not meet the preset conditions, and the set obtained from the last update is used as the initial evidence set.
[0169] Optionally, the incremental benefit is quantified by the difference in the objective function value before and after adding a candidate piece of evidence; the objective function includes a first calculation term, a second calculation term, and a third calculation term; the first calculation term is used to represent the explanatory power of the evidence set for the target anomalous event, the second calculation term is used to represent the degree of conflict among candidate evidence within the evidence set, and the third calculation term is used to represent the stable support power of the evidence set for the root cause conclusion.
[0170] Optionally, the incremental revenue calculation unit is specifically used to: calculate the second calculation term for the target set using the following calculation method: Calculate the temporal conflict degree, topological conflict degree, and statistical conflict degree for each candidate piece of evidence within the target set; The overall conflict degree of the candidate evidence is calculated based on at least two of the temporal conflict degree, topological conflict degree, and statistical conflict degree corresponding to the same candidate evidence. The total conflict degree of the target set is calculated based on the comprehensive conflict degree of each candidate piece of evidence within the target set, and is used as the second calculation term.
[0171] Optionally, the incremental revenue calculation unit is specifically used to calculate the first calculation term for the target set using the following calculation method: Based on the evidence type of each candidate piece of evidence in the target set and the explanatory power measurement method corresponding to the evidence type, determine the basic explanatory power corresponding to each candidate piece of evidence. Determine the pruned subset corresponding to each candidate piece of evidence within the target set; the pruned subset lacks corresponding candidate evidence compared to the target set; Obtain the stability metrics corresponding to each subset of the deleted versions and the target set respectively; For each candidate piece of evidence within the target set, a stability factor is calculated based on the stability index of the corresponding pruned subset and the stability index of the target set. The overall explanatory power of the candidate evidence is calculated based on the basic explanatory power and stability factor corresponding to the same candidate evidence. The sum of the comprehensive explanatory power of the target set is calculated based on the comprehensive explanatory power of each candidate piece of evidence within the target set and is used as the first calculation item.
[0172] Optionally, the counterfactual pruning module 43 includes: The indicator determination unit is used to determine the stability indicator and objective function value corresponding to the counterfactual evidence set after replaying the counterfactual evidence set; the objective function includes a first calculation term, a second calculation term, and a third calculation term; the first calculation term is used to represent the explanatory power of the evidence set for the target anomalous event, the second calculation term is used to represent the degree of conflict of candidate evidence within the evidence set, and the third calculation term is used to represent the stable support power of the evidence set for the root cause conclusion; The set judgment unit is used to determine that the counterfactual evidence set satisfies the preset stability requirement and the preset explanatory power requirement when the stability index is not lower than the preset stability threshold and the objective function value is not lower than the preset explanatory power threshold.
[0173] The evidence set acquisition module 41 is specifically used for: The observation time window is determined based on the alarm time of the target abnormal event. Evidence corresponding to the data whose time range intersects with the observation time window is extracted from the multi-source data to form a time-based preliminary screening evidence set. Based on the set of alarm entities related to the target abnormal event and the service and resource topology graph, determine the set of alarm context nodes; Evidence from the time-based initial screening evidence set that the source entity is located within the alarm context node set is selected to form a context-related evidence set. The candidate evidence set is obtained by removing evidence that does not meet the requirements in terms of explanatory power or data source quality, as well as duplicate evidence, from the context-related evidence set.
[0174] Optionally, the evidence screening device for root cause localization also includes: an evidence chain construction module; The evidence chain construction module is specifically used for: Based on the target evidence set, the abnormal related nodes covered by the evidence are determined, and the connected subgraphs corresponding to the abnormal related nodes are extracted from the service and resource topology graph. Based on the time information corresponding to each piece of evidence in the target evidence set, the propagation order of the abnormal related nodes in the connected subgraph is determined; Search the root cause candidate node for the propagation path to the affected node in the connected subgraph, and select representative evidence corresponding to the path node along the propagation path. Construct a minimum chain of evidence based on the propagation path, the propagation order, and the representative evidence.
[0175] It should be noted that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments. The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separate. The components indicated as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment solution according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0176] The above description is merely one specific implementation of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for screening evidence oriented towards root cause localization, characterized in that, include: Acquire multi-source data related to the target anomalous event, and determine a set of candidate evidence based on the multi-source data; Based on the explanatory power of the candidate evidence in the candidate evidence set for the target anomalous event, the degree of conflict between the candidate evidence, and the stable support of the candidate evidence set for the root cause conclusion, evidence is screened from the candidate evidence set to obtain an initial evidence set. For each candidate piece of evidence in the initial evidence set, a counterfactual evidence set is constructed by deleting the corresponding candidate piece of evidence, and the counterfactual evidence set is determined by replay to see if it meets the preset stability requirements and preset explanatory power requirements. If the counterfactual evidence set satisfies the preset stability requirement and the preset explanatory power requirement, the corresponding candidate evidence is deleted from the initial evidence set. Repeat the candidate evidence deletion judgment until there is no candidate evidence that still meets the preset stability requirement and the preset explanatory power requirement after deletion, and obtain the target evidence set.
2. The evidence triage method for root cause positioning according to claim 1, wherein, The initial evidence set is obtained by screening evidence from the candidate evidence set based on the explanatory power of the candidate evidence in the candidate evidence set for the target anomalous event, the degree of conflict between the candidate evidence, and the stable support of the candidate evidence set for the root cause conclusion, including: Establish an empty current evidence set; For each candidate piece of evidence in the candidate evidence set, calculate the incremental benefit formed by adding the corresponding candidate piece of evidence to the current evidence set; the incremental benefit is used to characterize the impact of adding the corresponding candidate piece of evidence, and the impact includes: the impact on the explanatory power of the current evidence set, the impact on the degree of conflict between the evidence in the current evidence set, and the impact on the stability support capability of the current evidence set; Based on the incremental gains, target candidate evidence is determined and added to the current evidence set; Based on the remaining candidate evidence after the target candidate evidence is added to the current evidence set and the updated current evidence set, the incremental benefit corresponding to each remaining candidate evidence is calculated, and the candidate evidence selection is repeated to continuously update the current evidence set. The set obtained from the last update is used as the initial evidence set.
3. The evidence triage method for root cause positioning of claim 2, wherein, The step of calculating the incremental benefit for each candidate piece of evidence in the candidate evidence set after adding the corresponding candidate piece of evidence to the current evidence set specifically includes: For each candidate piece of evidence in the candidate evidence set, determine whether the updated current evidence set satisfies the preset constraints after adding the corresponding candidate piece of evidence to the current evidence set. For the updated current evidence set that satisfies the preset constraints, calculate the resulting incremental revenue respectively; The process involves adding the target candidate evidence to the current evidence set, then using the remaining candidate evidence and the updated current evidence set as a basis. This process continues to calculate the incremental benefit for each remaining candidate evidence, and the candidate evidence selection is repeated to continuously update the current evidence set. The final updated set is used as the initial evidence set. Specifically, this includes: Based on the remaining candidate evidence after the target candidate evidence is added to the current evidence set and the updated current evidence set, the incremental benefits corresponding to each remaining candidate evidence are calculated under the preset constraints, and the candidate evidence selection is repeated to continuously update the current evidence set until there are no remaining candidate evidences in the candidate evidence set that satisfy the preset constraints, or until the incremental benefits corresponding to the remaining candidate evidence do not satisfy the preset conditions. The set obtained from the last update is then used as the initial evidence set.
4. The evidence filtering method for root cause oriented positioning according to claim 2, characterized in that, The incremental benefit is quantified by the difference in the objective function value before and after adding a candidate piece of evidence; the objective function includes a first calculation term, a second calculation term, and a third calculation term; the first calculation term is used to represent the explanatory power of the evidence set for the target anomalous event, the second calculation term is used to represent the degree of conflict among candidate evidence within the evidence set, and the third calculation term is used to represent the stable support power of the evidence set for the root cause conclusion.
5. The evidence triage method for root cause positioning of claim 4, wherein, For the target set, the calculation method for the second calculation item includes: Calculate the temporal conflict degree, topological conflict degree, and statistical conflict degree for each candidate piece of evidence within the target set; The overall conflict degree of the candidate evidence is calculated based on at least two of the temporal conflict degree, topological conflict degree, and statistical conflict degree corresponding to the same candidate evidence. The total conflict degree of the target set is calculated based on the comprehensive conflict degree of each candidate piece of evidence within the target set, and is used as the second calculation term.
6. The evidence triage method for root cause positioning of claim 4, wherein, For the target set, the calculation method for the first calculation item includes: Based on the evidence type of each candidate piece of evidence in the target set and the explanatory power measurement method corresponding to the evidence type, determine the basic explanatory power corresponding to each candidate piece of evidence. Determine the pruned subset corresponding to each candidate piece of evidence within the target set; the pruned subset lacks corresponding candidate evidence compared to the target set; Obtain the stability metrics corresponding to each subset of the deleted versions and the target set respectively; For each candidate piece of evidence within the target set, a stability factor is calculated based on the stability index of the corresponding pruned subset and the stability index of the target set. The overall explanatory power of the candidate evidence is calculated based on the basic explanatory power and stability factor corresponding to the same candidate evidence. The sum of the comprehensive explanatory power of the target set is calculated based on the comprehensive explanatory power of each candidate piece of evidence within the target set and is used as the first calculation item.
7. The evidence triage method for root cause positioning of claim 1, wherein, The step of determining whether the counterfactual evidence set meets the preset stability requirements and preset explanatory power requirements through playback includes: After replaying the counterfactual evidence set, the stability index and objective function value corresponding to the counterfactual evidence set are determined; the objective function includes a first calculation term, a second calculation term, and a third calculation term; the first calculation term is used to represent the explanatory power of the evidence set for the target anomalous event, the second calculation term is used to represent the degree of conflict of candidate evidence within the evidence set, and the third calculation term is used to represent the stable support power of the evidence set for the root cause conclusion; When the stability index is not lower than a preset stability threshold and the objective function value is not lower than a preset explanatory power threshold, the counterfactual evidence set is determined to meet the preset stability requirement and the preset explanatory power requirement.
8. The evidence triage method for root cause positioning of claim 1, wherein, The determination of the candidate evidence set based on the multi-source data includes: The observation time window is determined based on the alarm time of the target abnormal event. Evidence corresponding to the data whose time range intersects with the observation time window is extracted from the multi-source data to form a time-based preliminary screening evidence set. Based on the set of alarm entities related to the target abnormal event and the service and resource topology graph, determine the set of alarm context nodes; Evidence from the time-based initial screening evidence set that the source entity is located within the alarm context node set is selected to form a context-related evidence set. The candidate evidence set is obtained by removing evidence that does not meet the requirements in terms of explanatory power or data source quality, as well as duplicate evidence, from the context-related evidence set.
9. The evidence triage method for root cause positioning of claim 1, wherein, After obtaining the target set of evidence, the method further includes: Based on the target evidence set, the abnormal related nodes covered by the evidence are determined, and the connected subgraphs corresponding to the abnormal related nodes are extracted from the service and resource topology graph. Based on the time information corresponding to each piece of evidence in the target evidence set, the propagation order of the abnormal related nodes in the connected subgraph is determined; Search the root cause candidate node for the propagation path to the affected node in the connected subgraph, and select representative evidence corresponding to the path node along the propagation path. Construct a minimum chain of evidence based on the propagation path, the propagation order, and the representative evidence.
10. An evidence screening device for root cause localization, characterized in that, include: The evidence set acquisition module is used to acquire multi-source data related to the target abnormal event and determine a candidate evidence set based on the multi-source data. The first screening module is used to screen evidence from the candidate evidence set based on the explanatory power of the candidate evidence in the candidate evidence set for the target abnormal event, the degree of conflict between the candidate evidence, and the stable support power of the candidate evidence set for the root cause conclusion, to obtain an initial evidence set. The counterfactual pruning module is used to construct a counterfactual evidence set after deleting the corresponding candidate evidence for each candidate evidence in the initial evidence set, and to determine whether the counterfactual evidence set meets the preset stability requirements and preset explanatory power requirements by replaying the evidence; if the counterfactual evidence set meets the preset stability requirements and preset explanatory power requirements, the corresponding candidate evidence is deleted from the initial evidence set. Repeat the candidate evidence deletion judgment until there is no candidate evidence that still meets the preset stability requirement and the preset explanatory power requirement after deletion, and obtain the target evidence set.