Multi-dimensional data root cause positioning method and device based on reconstruction and heuristic search

By adopting reconstruction and heuristic search methods in the microservice architecture, using the self-coding neural network and an improved potential scoring mechanism, the problem of detecting subtle anomalies in multidimensional data in the microservice architecture is solved, and the rapid and accurate root cause positioning is achieved to adapt to the imbalance of long-tail distribution of multidimensional data.

CN120256175APending Publication Date: 2025-07-04XIAN UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510243638.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In microservice architecture, existing methods are difficult to quickly and accurately detect and locate subtle but critical anomalies in multidimensional data, resulting in potential problems not being discovered in time. When dealing with imbalance in the long-tail distribution of multidimensional data, the root cause is insufficient accuracy and stability of positioning.

Method used

The multidimensional data root cause positioning method based on reconstruction and heuristic search is adopted. By obtaining the observations of multidimensional data, the anomaly degree value is calculated, the data is reconstructed using the self-encoded neural network, and combined with the generalized ripple effect and improved potential scoring mechanism, the screening of significant anomalies and subtle key anomalies and root cause positioning is carried out.

Benefits of technology

It realizes rapid and timely discovery of potential problems, improves the accuracy of positioning of multi-dimensional data root causes, can effectively deal with extreme value interference when long-tail distribution is uneven, focuses on the fault area, reduces the false alarm rate and improves the detection sensitivity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256175A_ABST
    Figure CN120256175A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-dimensional data root cause positioning method and device based on reconstruction and heuristic search, and the method comprises the steps: obtaining an observation value of multi-dimensional data, carrying out the data prediction according to the observation value, and obtaining a prediction value of the multi-dimensional data; calculating an abnormal degree value of each leaf node dimension combination according to the observed value and the predicted value; screening significant abnormal data and non-significant abnormal data from the multi-dimensional data according to the abnormal degree value; performing data reconstruction on the data without significant anomalies by using a pre-trained self-encoding neural network, and calculating a reconstruction error; according to the reconstruction error, continuing to screen fine key abnormal data from the data without significant anomalies; and carrying out multi-dimensional data root cause positioning by utilizing a heuristic search mechanism according to the significant abnormal data and the subtle key abnormal data. According to the method, potential problems can be quickly and timely found, interference of extreme values when multidimensional data long tail distribution is unbalanced can be effectively dealt with, the method focuses on a fault area, and the positioning precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of information technology, and particularly relates to a multi-dimensional data root cause localization method and device based on reconstruction and heuristic search. Background Art

[0002] With the rapid development of information technology, the microservices architecture has gradually become an important trend in modern software development (such as mobile banking, e-commerce, online education, etc.). Compared with traditional monolithic applications, the microservices architecture splits an application into multiple small and independent services, greatly improving the maintainability, scalability, and flexibility of the system. However, with the expansion of the system scale and the increase in the complexity of interactions between services, the stability of the system faces severe challenges. Despite investing a large amount of resources in round-the-clock monitoring of software service maintenance, failures are still inevitable. The highly distributed nature of the microservices architecture means that any minor anomaly may trigger a chain reaction, resulting in anomalies in multiple services and even causing a system-level crash. Therefore, accurately detecting anomalies and locating the root cause of failures is crucial for ensuring the stable operation of the system and optimizing the user experience.

[0003] In a microservices system, root cause localization refers to accurately identifying the root cause that triggers system anomalies from multi-dimensional attributes by using anomaly alarm information and key performance indicators. Common failure scenarios include abnormal traffic patterns, packet loss, and traffic fluctuations. The diversification of system failure forms poses a huge challenge to the root cause localization of microservices anomalies. As the number of attributes or the number of values for each attribute increases, the search space will grow exponentially. Therefore, operation and maintenance personnel usually monitor key performance indicators (KPIs), which can effectively reflect the state and performance of the system. When a failure occurs, the measured values of a specific combination of attributes will show anomalies, becoming the primary signal of potential failures. This combination of attributes can effectively indicate the location of the failure and provide key clues for root cause localization. By identifying these specific combinations of attributes, operation and maintenance personnel can locate the source of the failure, and these combinations of attributes are called the roots of multi-dimensional data.

[0004] In a microservices architecture, the root cause localization of abnormal indicators is usually divided into two stages: anomaly detection of performance indicators and root cause localization of abnormal indicators. The main limitations of existing methods for root cause localization in a microservices architecture are at least reflected in the following aspects: First, the complex call relationships between services make it challenging to efficiently retrieve significant anomalies in a large amount of indicator data with a low anomaly rate; Second, existing methods mostly focus on significant anomalies and perform poorly in detecting small but critical anomalies, resulting in potential problems not being discovered in a timely manner. Summary of the Invention

[0005] In order to solve the above problems existing in the prior art, the present invention provides a multi-dimensional data root cause localization method and device based on reconstruction and heuristic search.

[0006] The technical problem to be solved by the present invention is achieved through the following technical solutions:

[0007] A multi-dimensional data root cause localization method based on reconstruction and heuristic search, comprising:

[0008] Obtain the observed values of multi-dimensional data, and perform data prediction based on the observed values to obtain the predicted values of the multi-dimensional data;

[0009] According to the observed values and the predicted values, calculate the abnormality degree values of each leaf node dimension combination; the leaf node dimension combination is the dimension combination corresponding to the finest-grained data in the multi-dimensional data;

[0010] According to the abnormality degree values, screen out significant abnormal data and non-significant abnormal data from the multi-dimensional data;

[0011] Use a pre-trained auto-encoder neural network to reconstruct the non-significant abnormal data and calculate the reconstruction error;

[0012] According to the reconstruction error, continue to screen out subtle key abnormal data from the non-significant abnormal data; the abnormality of the subtle key abnormal data is more subtle than that of the significant abnormal data and belongs to key abnormality;

[0013] According to the significant abnormal data and the subtle key abnormal data, use a heuristic search mechanism to perform multi-dimensional data root cause localization.

[0014] Optionally, calculating the abnormality degree value of each leaf node dimension combination according to the observed value and the predicted value includes:

[0015]

[0016] Wherein, v d (ε) represents the observed value of the leaf node dimension combination ε for which the abnormality degree value is currently to be calculated, and f d (ε) represents the predicted value of this leaf node dimension combination ε; v d (ε i ) represents the observed values of other leaf node dimension combinations ε i at the same moment, and f d (ε i ) represents the predicted values of other leaf node dimension combinations ε i , n is the number of other leaf node dimension combinations at the same moment, and AD(ε) is the abnormality degree value of the leaf node dimension combination ε.

[0017] Optionally, the observed values of the multi-dimensional data include the observed values at multiple time points;

[0018] Performing data prediction based on the observed values to obtain predicted values of multi-dimensional data, including:

[0019] Using a sliding time window to segment the observed values at the multiple time points;

[0020] Taking the average of the observed values at the time points within each sliding time window to obtain the predicted value at each time point within that sliding time window.

[0021] Optionally, screening significant abnormal data and non-significant abnormal data from the multi-dimensional data according to the abnormal degree value, including:

[0022] Obtaining the CDF curve of the abnormal degree value;

[0023] Determining a first screening threshold using the knee point method according to the CDF curve;

[0024] Screening significant abnormal data and non-significant abnormal data from the multi-dimensional data using the first screening threshold.

[0025] Optionally, the pre-trained autoencoder neural network includes an encoder and a decoder;

[0026] Reconstructing the non-significant abnormal data using the pre-trained autoencoder neural network, specifically by inputting the non-significant abnormal data into the pre-trained autoencoder neural network to achieve the following operations:

[0027] The encoder compresses the non-significant abnormal data through a linear transformation to extract the low-dimensional latent feature representation of the non-significant abnormal data:

[0028] z = f θ (x) = σ(W e x + b e );

[0029] where z represents the low-dimensional latent feature representation, represents the non-significant abnormal data of dimension d in the real number field , f θ (·) represents the encoder, is the weight matrix of the encoder, k << d, is the bias vector of the encoder;

[0030] The decoder reconstructs the non-significant abnormal data according to the low-dimensional latent feature representation:

[0031]

[0032] where, represents the reconstructed data, g θ'(·) represents the decoder, σ(·) is the non-linear activation function Sigmoid, is the weight matrix of the decoder, is the bias vector of the decoder.

[0033] Optionally, continuing to screen out subtle critical abnormal data from the non-significant abnormal data according to the reconstruction error includes:

[0034] Calculating the mean absolute error MAE of the reconstruction error;

[0035] Modeling the distribution of the mean absolute error MAE by using the kernel density estimation method to obtain the MAE distribution;

[0036] Determining the second screening threshold according to the MAE distribution by using the knee point method;

[0037] Continuing to screen out subtle critical abnormal data from the non-significant abnormal data by using the second screening threshold.

[0038] Optionally, performing multi-dimensional data root cause location by using a heuristic search mechanism according to the significant abnormal data and the subtle critical abnormal data includes:

[0039] Combining the significant abnormal data and the subtle critical abnormal data into an abnormal data set;

[0040] Clustering the abnormal data set through the generalized ripple effect to obtain multiple clustering clusters;

[0041] For each clustering cluster, calculating the descending scores of all non-leaf node dimension combinations within the clustering cluster, screening multiple first candidate non-leaf node dimension combinations according to the descending scores; calculating the internal influence degree and the cluster influence degree of each first candidate non-leaf node dimension combination; if the internal influence degree and the cluster influence degree of a first candidate non-leaf node dimension combination are both higher than the corresponding thresholds, then taking it as a second candidate non-leaf node dimension combination and calculating its potential score; generating a comprehensive evaluation index according to the internal influence degree, the cluster influence degree and the potential score of each second candidate non-leaf node dimension combination; finding the root cause dimension combination of the clustering cluster from all the second candidate non-leaf node dimension combinations of the clustering cluster; wherein, the potential score is used to evaluate the abnormal degree of the non-leaf node dimension combination by using relative difference.

[0042] Optionally, the calculation method of the potential score is as follows:

[0043]

[0044] Among them, v(e1) represents the observed value of the leaf node dimension combination under the second candidate non-leaf node dimension combination for which the potential score is to be calculated, a(e1) represents the expected value of the leaf node dimension combination under the second candidate non-leaf node dimension combination, f(e1) represents the observed values of the other leaf node dimension combinations in the cluster where the second candidate non-leaf node dimension combination is located, f(e2) represents the predicted values of the other leaf node dimension combinations in the cluster where the second candidate non-leaf node dimension combination is located, avg(·) represents taking the average of the calculation results in (·) among the leaf node dimension combinations, and IPS represents the potential score.

[0045] The present invention also provides a multi-dimensional data root cause location device based on reconstruction and heuristic search, including:

[0046] An acquisition module, configured to acquire the observed values of multi-dimensional data and perform data prediction based on the observed values to obtain the predicted values of the multi-dimensional data;

[0047] A calculation module, configured to calculate the abnormality degree value of each leaf node dimension combination according to the observed values and the predicted values; the leaf node dimension combination is the dimension combination corresponding to the finest-grained data in the multi-dimensional data;

[0048] A first screening module, configured to screen out significantly abnormal data and non-significantly abnormal data from the multi-dimensional data according to the abnormality degree value;

[0049] A reconstruction module, configured to use a pre-trained autoencoder neural network to reconstruct the non-significantly abnormal data and calculate the reconstruction error;

[0050] A second screening module, configured to further screen out subtle key abnormal data from the non-significantly abnormal data according to the reconstruction error; the abnormality of the subtle key abnormal data is more subtle than that of the significantly abnormal data and belongs to key abnormality;

[0051] A location module, configured to perform multi-dimensional data root cause location using a heuristic search mechanism according to the significantly abnormal data and the subtle key abnormal data.

[0052] The present invention also provides an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus.

[0053] The memory is used to store a computer program;

[0054] The processor, when executing the computer program stored on the memory, implements any one of the above multi-dimensional data root cause location methods based on reconstruction and heuristic search.

[0055] The multi-dimensional data root cause localization method based on reconstruction and heuristic search provided by the present invention uses the anomaly degree value to quickly screen out significant anomalies, and then reconstructs normal data through a pre-trained autoencoder neural network. Based on the reconstruction error, subtle key anomaly data is further screened to capture subtle but key anomalies. Thus, according to the significant anomaly data and the subtle key anomaly data, a heuristic search mechanism is used for multi-dimensional data root cause localization, which can quickly and timely discover potential problems.

[0056] In addition, the present invention proposes an improved potential score (IPS) based on the generalized ripple effect (GRE), which combines the clustering filtering (CF) and leaf filtering (LF) strategies to effectively cope with the interference of extreme values when the long-tail distribution of multi-dimensional data is unbalanced, focus on the fault area and improve the localization accuracy.

[0057] The following will further elaborate on the present invention in conjunction with the accompanying drawings. Description of the Drawings

[0058] Figure 1 is a schematic flowchart of a multi-dimensional data root cause localization method based on reconstruction and heuristic search provided by an embodiment of the present invention;

[0059] Figure 2 shows a cuboid with 3 attributes;

[0060] Figure 3 shows the changing trend of AD(ε) with the increase of λ proposed in the present invention;

[0061] Figure 4 shows the working process of the present invention;

[0062] Figure 5 shows the performance of the present invention and several other existing algorithms in six different datasets under the conditions of different numbers of elements n and cuboid layer settings;

[0063] Figure 6 shows the F1-score of the present invention and several other existing algorithms for root cause localization only based on the anomaly degree;

[0064] Figure 7 shows the F1-score of the present invention and several other existing algorithms for root cause localization only based on the data reconstruction error;

[0065] Figure 8 shows the F1-score of the present invention and several other existing algorithms for root cause localization only based on IPS;

[0066] Figure 9 shows the F1-score of the present invention and several other existing algorithms for root cause localization based on the anomaly degree and the data reconstruction error;

[0067] Figure 10 Shows the F1-scores of the present invention and several other existing algorithms for root cause location based on the degree of anomaly and IPS;

[0068] Figure 11 Shows the F1-scores of the present invention and several other existing algorithms for root cause location based on data reconstruction error and IPS;

[0069] Figure 12 Shows the comparison results of the single fault running time of the present invention and several other existing algorithms on the A dataset;

[0070] Figure 13 Shows the comparison results of the single fault running time of the present invention and several other existing algorithms on the D dataset;

[0071] Figure 14 Shows the comparison results of the single fault running time of the present invention and several other existing algorithms on the B5 dataset;

[0072] Figure 15 Shows the comparison results of the single fault running time of the present invention and several other existing algorithms on the B6 dataset;

[0073] Figure 16 Shows the comparison results of the single fault running time of the present invention and several other existing algorithms on the B7 dataset;

[0074] Figure 17 Shows the comparison results of the single fault running time of the present invention and several other existing algorithms on the B8 dataset;

[0075] Figure 18 Shows the F1-scores of the present invention when the number of cuboid layers = 3 and n_element = 3 under different δ. Detailed implementation manners

[0076] The present invention will be further described in detail below with reference to specific embodiments, but the implementation manners of the present invention are not limited thereto.

[0077] In a microservices system, root cause localization is generally divided into anomaly detection of performance metrics and root cause localization of anomalous metrics. The accuracy of pre - anomaly detection is crucial for the effectiveness of root cause localization. Existing anomaly detection techniques are widely applied to univariate and multivariate time - series data analysis. Univariate mainly focuses on time - series data from a single source, while in a multivariate scenario, the interaction between multiple time - series data is of great significance for accurate anomaly detection. Traditional methods, such as random forest, support vector machine, and K - Means clustering, identify anomaly points by modeling the time - series distribution. Some studies use the prediction residual as an indicator to measure the degree of change in the combination of attributes, and determine whether an anomaly occurs by setting a threshold. The prediction residual mainly reflects the overall performance of the model, but it is difficult to identify subtle and critical anomalies in multi - dimensional data, which is prone to false alarms, reducing the detection accuracy and weakening the early warning ability of the system. In addition, anomaly patterns often involve multi - dimensional correlation and mutation relationships, and it is difficult to comprehensively identify complex anomalies in multi - dimensional interactions by relying solely on the prediction residual.

[0078] In view of the above technical status quo, in order to quickly and timely discover potential problems, the embodiments of the present invention provide a multi - dimensional data root cause localization method based on reconstruction and heuristic search. This method can not only quickly screen out significant anomalies in a large amount of operation and maintenance data, but also detect ignored subtle or complex anomalies. Refer to Figure 1 and in combination with Figure 4 as shown, the method includes the following steps:

[0079] S10. Obtain the observed values of multi - dimensional data, and perform data prediction based on the observed values to obtain the predicted values of the multi - dimensional data.

[0080] In a microservices system, the recording of running state information is crucial for root cause localization. Metrics are standards used to quantify and monitor the system state, and are generally divided into two categories: service node metrics (such as CPU utilization) and inter - service call metrics (such as request time), and the latter reflects the state of service calls. In order to effectively monitor and ensure the reliability of the online service system, Internet enterprises generally monitor a set of key performance indicators (KPIs) such as response time, memory usage, network transmission rate, etc., to continuously track the system state and discover potential relationships between multi - variable data. Table 1 schematically shows an example of multi - dimensional data at a specific time point, which is multi - dimensional KPI (key performance indicator) data with three dimensions (province, network service provider, website), and records all the finest - grained data every 5 minutes. Each record contains the KPI for each time interval of the attribute values, such as (Beijing; China Unicom; Weibo). The finest - grained KPI records can be summarized into coarser - grained KPIs, such as (Beijing; China Unicom; *), where * is a wildcard, and the wildcard indicates that the value of this dimension can be any possible value. Therefore, the wildcard dimension is not fixed, but can match multiple different dimension values.

[0081] In multi-dimensional data, each dimension (such as province, operator, website, etc.) represents an attribute or category, and its value is the specific value of this attribute. A dimension combination refers to a specific combination of values of these dimensions, which is equivalent to a definite point in multi-dimensional space. For example, the dimension combination {Beijing, China Unicom, Weibo} means that the province is "Beijing", the operator is "China Unicom", and the website is "Weibo" under this dimension. Figure 2 A cuboid with 3 attributes is shown. A cuboid is an enumeration of a set of attribute combinations, containing all leaf nodes that share the same non-wildcard dimension combination. For example, the cuboid {*, China Mobile, Weibo} represents the combination of leaf nodes in different provinces when the operator is China Mobile and the website is Weibo. A leaf node is the most specific dimension combination without any wildcards, representing a complete dimension combination. {Beijing, China Mobile, Weibo} and {Shanghai, China Mobile, Weibo} are both leaf nodes, and they belong to the cuboid {*, China Mobile, Weibo}. For a scenario with d dimensions, there are a total of 2 d -1 cuboids. The actual value of a dimension combination is obtained by aggregating the original logs and calculating the corresponding actual KPI value, while the predicted value of the dimension combination can be deduced from historical data. Table 1 schematically shows Figure 2 the actual value and predicted value of each leaf node dimension combination in

[0082] Table 1 Multi-dimensional root cause cases

[0083]

[0084]

[0085] When the microservice system detects a potential fault (usually triggered by an alarm from the monitoring system), it will automatically start, and use the multi-dimensional observed values and their predicted values at the moment of the fault as the input of the method proposed in the present invention to deeply analyze the data and locate the root cause of the fault.

[0086] In a complex online service system, the anomaly detection of KPIs and the scientific evaluation of their severity are crucial for ensuring system stability and accurate fault location. Facing the characteristics of massive operation and maintenance data and a low anomaly rate, it is a huge challenge to detect anomalies quickly and accurately. Therefore, the present invention designs an anomaly detection mechanism based on multi-variate time series data.

[0087] Specifically, the observed values of multi-dimensional data include the observed values at multiple time points, expressed as a set of multi-variate time series data X = {X1, X2,..., X n} with a length of n, where the observed value at each time point t (1 ≤ t ≤ n) is an m-dimensional vector X t= [x1, x2,..., x m represents the system state information at this time.

[0088] Correspondingly, based on the observed values of the multi-dimensional data, data prediction is performed to obtain the predicted values of the multi-dimensional data, including:

[0089] (1-1) Use a sliding time window to segment the observed values at multiple time points;

[0090] For example, to capture the dynamic change patterns in the time series, a sliding time window of length k is used to segment the time series data X, generating a set of subsequences W = {W1, W2,..., W j}, where the subsequence W t = {X t-k+1 , X t-k+2 ,..., X t} represents the sliding window data at time t, j is the total number of subsequences, and subsequent anomaly detection of the data X t can be transformed into anomaly detection of the subsequence W t .

[0091] (1-2) Take the average of the observed values at the time points within each sliding time window to obtain the predicted values at each time point within the sliding time window.

[0092] Here, obtaining the predicted value by taking the average is a specific implementation method of the time series prediction algorithm. In practice, there are multiple specific implementation methods for the time series prediction algorithm, and other optional implementation methods can also be used for data prediction in the present invention, such as the moving average (MA) algorithm, etc.

[0093] In addition, the sliding window technology can be combined with time series analysis to construct a performance baseline (PB) for the service components (functional modules) of each microservice system. The PB represents the mean or trend of KPIs within a period of time before an anomaly occurs. Based on this performance baseline for data prediction can effectively suppress the interference of short-term fluctuations on the prediction results.

[0094] S20. Calculate the anomaly degree value of each leaf node dimension combination according to the observed values and predicted values of the multi-dimensional data.

[0095] Here, the leaf node dimension combination is the dimension combination corresponding to the finest-grained data in the multi-dimensional data. For example Figure 2 the cuboid corresponding to the third layer in

[0096] corresponds to a leaf node dimension combination. In this step S20, the anomaly degree value of each leaf node dimension combination can be calculated using the GRE (Generalized Ripple Effect) theory, but is not limited thereto.

[0097] Specifically, the ripple effect is used to describe the relationship between the abnormal amplitude changes of different combinations of attributes caused by the same root cause. The generalized ripple effect can capture the fault propagation and mutual influence in the system caused by the same root cause. It is applicable to processing multi-dimensional data and different types of metrics. The GRE indicates that combinations of attributes affected by the same root cause will change in the same proportion, and it is applicable to both additive and non-additive KPIs. For example, most common derived metrics (such as success rate, etc.) still conform to the basic principle of the ripple effect. According to the GRE, if a root cause S and its derived leaf node e exist, the following relational expression should be satisfied:

[0098]

[0099] where f(·) is the predicted value, v(·) is the actual value, and S is the root cause.

[0100] However, in some actual scenarios, the predicted value may be zero, and then formula (1) is no longer applicable. To avoid the problem of a zero denominator, f is replaced with Then formula (1) can be rewritten as:

[0101]

[0102] Formula (2) is consistent with formula (1) in the following aspects:

[0103] 1. Invalid data processing: When f(e) = v(e) = 0 or f(S) = v(S) = 0, the relationship between e and S in the formula loses its meaning, and the formula has no practical significance at this time.

[0104] 2. Consistency: If f(e) ≠ 0 and f(S) ≠ 0, formula (2) is equivalent to formula (1), ensuring the inheritance of the original method.

[0105] 3. Special case handling: When f(e) = 0 ≠ v(e) or f(S) = 0 ≠ v(S), formula (1) is no longer applicable, while formula (2) can effectively handle these boundary cases by adding which can effectively handle these boundary cases, expand the application scope of the formula, and improve its robustness.

[0106] The deviation score d(e) is a key metric in the GRE and is defined as follows:

[0107]

[0108] According to the GRE principle, leaf nodes caused by the same root cause have similar deviation scores, while leaf nodes caused by different root causes have different deviation scores. Therefore, the deviation score can be used to cluster leaf nodes (leaf node dimension combinations) into similar clusters. For example, the deviation scores of the root causes {*, China Unicom, Weibo} in Table 1 are as follows:

[0109]

[0110] The deviation scores of the corresponding leaf nodes in Table 1 are 0.33, 0.35, and 0.33, which conform to the GRE theory. Other leaf nodes that do not belong to this root cause show different deviation scores.

[0111] During the fault location process, assume that a non-leaf node S is a potential root cause node, and the leaf node e inherits from this node S. The abnormal leaf attribute combination should follow GRE. According to the GRE principle, the abnormal degree value of the leaf node dimension combination under the root cause S can be calculated by formula (4):

[0112]

[0113] Among them, f(e) represents the predicted value of the leaf node, and v(S) and f(S) represent the actual value and predicted value of the root cause node S respectively. The effectiveness of the root cause attribute combination can be determined by evaluating the difference |v(e)-a(e)| between its actual value and the expected abnormal value, where v(e) is the actual value of the leaf node.

[0114] The above a(e) only compares the KPI value of a single leaf combination with the deviation of its historical normal state, and cannot conduct a comprehensive analysis of the potential anomalies of the system. To solve this problem, in another implementation, according to the observed values and predicted values of multi-dimensional data, the abnormal degree value of each leaf node dimension combination can be calculated, which can be achieved through the following formula:

[0115]

[0116] where, v d (ε) represents the observed value of the leaf node dimension combination ε for which the abnormal degree value is to be calculated currently, f d (ε) represents the predicted value of this leaf node dimension combination ε; v d (ε i ) represents the observed values of other leaf node dimension combinations ε i at the same moment, f d (ε i ) represents the predicted values of other leaf node dimension combinations ε i at the same moment, n is the number of other leaf node dimension combinations at the same moment, and AD(ε) is the abnormal degree value of the leaf node dimension combination ε.

[0117] In this implementation, AD(ε) can accurately identify most obvious faults through double evaluation while ensuring the computational efficiency of the screening process.

[0118] Specifically, AD(ε) not only compares the deviation of the KPI value of a single blade combination from its historical normal state, but also makes a global comparison of its deviation with that of other blade combinations at the same moment, enhancing the comprehensive analysis of potential system anomalies. Let It can be found that AD(ε) ∈ (0, 1), Figure 3 which shows the changing trend of AD(ε) with the increase of λ. It can be seen that even when the deviation between the true measurement value v d (·) and the predicted value f d (·) is not significant, it can still effectively identify the anomaly.

[0119] In terms of computational complexity, |v d (ε) - f d (ε)| represents the absolute deviation between the true value v d (ε) and the predicted value f d (ε) at time t and dimension d, and the time complexity is O(1). The main computational overhead of formula (5) is concentrated in the cumulative calculation of the sum of deviations, that is, this item sums up the deviations at n - 1 moments. Since this step needs to traverse all samples and calculate the absolute deviation at each moment, its time complexity is O(n), where n represents the length of the time series data. The final result is a weighted deviation ratio, and the calculation steps include one addition and one multiplication operation. Therefore, the computational time complexity of formula (5) is O(n). In the scenario of large-scale data, this complexity improves the ability to retrieve significant anomalies in a large amount of data with a low anomaly rate, ensuring the real-time detection efficiency of the system.

[0120] S30. Screen significant anomaly data and non-significant anomaly data from the multi-dimensional data according to the anomaly degree values of the leaf node dimension combinations.

[0121] Specifically, screening significant anomaly data and non-significant anomaly data from the multi-dimensional data according to the anomaly degree values of the leaf node dimension combinations includes:

[0122] (3 - 1) Obtain the CDF (cumulative distribution function) curve of the anomaly degree values of the leaf node dimension combinations.

[0123] (3 - 2) Determine the first screening threshold using the knee point method according to the CDF curve.

[0124] Here, the knee point method is adopted for automatic threshold selection to accurately distinguish abnormal leaf attribute combinations from normal combinations. The knee point is defined as the point with the maximum curvature, which can effectively identify the abnormal combinations that only account for a small number among a large number of leaf attribute combinations. This method can determine the position of the threshold. Further increasing the threshold, the number of filtered leaf node dimension combinations will no longer decrease significantly.

[0125] (3-3) Filter significant abnormal data and non-significant abnormal data from the multi-dimensional data using the first screening threshold.

[0126] Specifically, using the first screening threshold, the data corresponding to the leaf node dimension combinations with abnormal degree values above and below this threshold are divided into two categories: significant abnormal data and non-significant abnormal data.

[0127] S40. Use the pre-trained autoencoder neural network to reconstruct the non-significant abnormal data and calculate the reconstruction error.

[0128] Although the abnormal degree value can quickly identify significant abnormal situations, it has limited ability in detecting subtle abnormalities. Such subtle abnormalities often show as small deviations in the initial stage, which may be early signals of potential system failures. However, these subtle abnormalities often hide key fault clues and are easily regarded as normal fluctuations and ignored. Therefore, the present invention further introduces an autoencoder neural network. By learning historical normal data, a normal behavior model of the system is constructed. By reconstructing the preliminarily screened normal data and based on the reconstruction error to identify abnormalities, the detection sensitivity is improved and false alarms are reduced.

[0129] The principle of the autoencoder detecting small but important abnormalities is that when the autoencoder processes normal data, its reconstruction error is usually small, while when encountering abnormal data, since these data deviate from the normal pattern, the reconstruction error will increase significantly. Therefore, the size of the reconstruction error becomes a key indicator for judging abnormalities.

[0130] Specifically, the pre-trained autoencoder neural network includes an encoder (Encoder) and a decoder (Decoder); using the pre-trained autoencoder neural network to reconstruct the non-significant abnormal data, specifically, inputting the non-significant abnormal data into the pre-trained autoencoder neural network to achieve the following operations:

[0131] (a) The encoder compresses the non-significant abnormal data through linear transformation and extracts the low-dimensional latent feature representation of the non-significant abnormal data:

[0132] z = f θ (x) = σ(W e x + b e )(6);

[0133] where z represents the low-dimensional latent feature representation, Denote in the real number field the data without significant anomalies with dimension d, f θ (·) represents the encoder, is the weight matrix of the encoder, k << d, is the bias vector of the encoder; the parameters of the encoder are represented as θ = {W e , b e}.

[0134] (b) The decoder reconstructs the data without significant anomalies based on the low-dimensional latent feature representation, that is, the decoder attempts to reconstruct data close to the original input from this low-dimensional latent feature representation:

[0135]

[0136] where, represents the reconstructed data, g θ' (·) represents the decoder, σ(·) is the non-linear activation function Sigmoid, is the weight matrix of the decoder, is the bias vector of the decoder, and the parameters of the decoder are represented as θ' = {W d , b d}.

[0137] The training objective of this autoencoder neural network is to minimize the reconstruction error between the original input x and the reconstructed output . The mean absolute error (MAE) can be used as the loss function, and its definition is:

[0138]

[0139] where, N is the dimension or the number of samples of the data input to the autoencoder neural network during training. MAE measures the absolute error between the input and the output during the data reconstruction process. By minimizing the MAE loss function, the autoencoder can effectively learn the latent structure of the input data and then optimize the model parameters.

[0140] S50. According to the reconstruction error, continue to screen out the subtle key abnormal data from the data without significant anomalies; the anomalies of this subtle key abnormal data are more subtle than those of the significant abnormal data and belong to key anomalies.

[0141] Specifically, according to the reconstruction error, continue to screen out the subtle key abnormal data from the data without significant anomalies, including:

[0142] (5-1) Calculate the mean absolute error MAE of the reconstruction error;

[0143] (5-2) Use the kernel density estimation method to model the distribution of the mean absolute error MAE to obtain the MAE distribution;

[0144] (5-3) Determine the second screening threshold according to the MAE distribution using the knee point method;

[0145] (5-4) Continue to screen out subtle key abnormal data from the data without significant anomalies using the second screening threshold.

[0146] Here, considering the complex and diverse abnormal forms in the microservices architecture, a single fixed reconstruction error threshold is difficult to adapt to all scenarios. The present invention combines kernel density estimation (KDE) to model the distribution of MAE and uses the knee point detection algorithm to automatically determine the abnormal threshold. For example, in the real case in Table 1, after reconstructing a set of sample data through an autoencoder model, the reconstruction error of each sample is calculated, and the second screening threshold is determined to be 1.46 through the knee point detection algorithm. Taking the first sample as an example, Real = 88, Predict = 94. The error value after reconstruction by the autoencoder is 0.495. Based on the standardized data, the reconstruction error of the model for this sample is 0.96, indicating that the prediction error of this sample is still within the normal range in the model reconstruction. On the contrary, the reconstruction error of sample 2 is 103.74, indicating that the predicted value of this sample has obvious anomalies during the reconstruction process. Therefore, this sample represents a potential abnormal point.

[0147] In a specific example, the present invention shows a two-stage anomaly detection algorithm for implementing the process of steps S10 to S50. The specific algorithm process is as follows:

[0148]

[0149] S60. Locate the root cause of multi-dimensional data using a heuristic search mechanism based on significant abnormal data and subtle key abnormal data.

[0150] After detecting an anomaly, the goal of root cause localization is to identify the specific cause of the anomaly. Existing methods include spectral analysis, random walk, causal inference, and heuristic search methods. Spectral analysis extracts system anomaly patterns through frequency-domain signals and is suitable for the localization of strong signals, but it is difficult to identify weak anomalies. The random walk algorithm realizes root cause localization by constructing a fault propagation graph and generating a probability transition matrix. However, the random walk depends on an accurate graph structure, and the processing complexity for high-order dependency relationships is relatively high. Causal inference methods infer root causes by establishing a multi-level causal relationship graph. However, the conditional independence detection of index data points results in a large computational overhead, affecting real-time performance. Compared with the above methods, the heuristic search method has gradually become a research hotspot because of its improved search efficiency in dealing with complex scenarios and its adaptability to the needs of different business scenarios. Among them, HotSpot proposes a root cause localization method based on RippleEffect, which quantifies the association between dimension combinations and leaf nodes through Potential Score, and effectively reduces the search space by combining Monte Carlo Tree Search (MCTS) and hierarchical pruning strategies. However, it is only applicable to additive KPI metrics and is difficult to handle more complex non-additive scenarios. Squeeze extends the Ripple Effect method based on HotSpot and solves the problems of derived metrics and zero-value prediction. Similar leaf nodes are clustered through the deviation score d(e), and the Generalized Potential Score (GPS) is used to improve generality and robustness. GPS evaluates the effectiveness of the root cause attribute combination by comparing the differences between the actual value, predicted value, and expected anomaly of the attribute combination. Its calculation formula is as follows:

[0151]

[0152] where S1 is the set of abnormal leaf attribute combinations of the root cause node; S2 is the set of all remaining normal leaf attribute combinations, and v, f, and a are the actual value, predicted value, and expected value of the leaf attribute combination, respectively. avg(·) represents taking the average of the calculation results in (·) among the dimension combinations of the leaf nodes.

[0153] However, when dealing with data with uneven long-tail distributions, relatively large differences may mask subtle but important changes, resulting in insufficient recognition of low-frequency anomalies. GPS cannot effectively reflect the extreme values in the low-frequency part in long-tail distribution data, affecting the accuracy and stability of root cause localization. In addition, when the actual value of a certain root cause node is large, traditional methods tend to underestimate its importance, resulting in misjudgment.

[0154] In view of the limitations of existing GPS methods in dealing with long-tailed distribution imbalance data, the present invention proposes IPS (Improved Potential Score), which uses relative difference to evaluate the abnormality degree of root cause nodes, assigns appropriate weights to low-frequency extreme values through a more flexible scoring mechanism to cope with the influence of extreme values in long-tailed distribution imbalance, ensures that even small-scale changes can be accurately identified, improves the accuracy of root cause location, and the IPS calculation formula is as follows:

[0155]

[0156] Among them, v(e1) represents the observed value of the leaf node dimension combination under the second candidate non-leaf node dimension combination for which the potential score is to be calculated, a(e1) represents the expected value of the leaf node dimension combination under the second candidate non-leaf node dimension combination, f(e1) represents the observed values of other leaf node dimension combinations in the clustering cluster where the second candidate non-leaf node dimension combination is located, f(e2) represents the predicted values of other leaf node dimension combinations in the clustering cluster where the second candidate non-leaf node dimension combination is located, avg(·) represents taking the average of the calculation results in (·) among leaf node dimension combinations, and IPS represents the potential score.

[0157] Based on the above IPS, in step S60, according to the significant abnormal data and subtle key abnormal data, a heuristic search mechanism is used for multi-dimensional data root cause location, including:

[0158] (1) Merge the significant abnormal data and subtle key abnormal data into an abnormal data set.

[0159] (2) Cluster the abnormal data set through the generalized ripple effect to obtain multiple clustering clusters.

[0160] Specifically, similar leaf nodes in the abnormal data set are clustered through the deviation score d(e) defined in the generalized ripple effect. In multi-dimensional fault location, each cuboid (i.e., dimension combination) covers all abnormal leaf attribute combinations within the clustering cluster, and some attribute subsets can exactly match the abnormal leaf combinations. According to the idea of the generalized ripple effect, if an attribute combination is the root cause attribute combination, then all its downstream leaf attribute combinations should show abnormalities, and these abnormal combinations should be clustered in the same abnormal cluster. Therefore, a necessary condition for the root cause attribute combination is that most of its downstream leaf attribute combinations are concentrated in the same abnormal cluster.

[0161] (3) For each clustering cluster, calculate the descent score r of all non-leaf node dimension combinations within the clustering cluster descended , and this descent score is used to measure the proportion of the downstream leaf combinations of the non-leaf node dimension combination (i.e., the leaf nodes with abnormal indicators among the leaf nodes it contains) in the clustering cluster; according to the descent score rdescended Screen multiple first candidate non-leaf node dimension combinations, specifically screen multiple first candidate non-leaf node dimension combinations with relatively high descending scores; calculate the internal influence degree and cluster influence degree of each first candidate non-leaf node dimension combination; if the internal influence degree and cluster influence degree of a first candidate non-leaf node dimension combination are both higher than the corresponding thresholds, then take it as a second candidate non-leaf node dimension combination, and calculate its potential score, which is used to evaluate the abnormality degree of the non-leaf node dimension combination by relative difference; generate a comprehensive evaluation index according to the internal influence degree, cluster influence degree and potential score of each second candidate non-leaf node dimension combination, for example, RS = avg(IPS + CF + LF) can be calculated as the comprehensive evaluation index, where CF represents the cluster influence degree and LF represents the internal influence degree; according to the comprehensive evaluation index, find the root cause dimension combination of this cluster from all second candidate non-leaf node dimension combinations of this cluster.

[0162] In the present invention, in the root cause location stage of the abnormal index, the significant abnormality and the key subtle abnormality results are combined, and a heuristic algorithm combining collaborative filtering and IPS is used to locate the real root cause of the failure. CF and LF, through a hierarchical filtering mechanism, while reducing the number of dimension combinations to be evaluated, ensure that the location process focuses on the most likely root cause feature combinations. IPS can effectively balance the influence of low-frequency extreme values on the location accuracy when dealing with the uneven long-tail distribution of multi-dimensional data, thus avoiding the false alarm problem caused by the abnormal distribution of data.

[0163] In practical applications, by comparing with GPS, IPS shows higher accuracy when dealing with long-tail distribution data. For example, in the real case in Table 1, the root cause is {*, ChinaUnicom, weibo.com}, and the scoring results of GPS and IPS are calculated as follows:

[0164]

[0165] The above calculation results can show that IPS can more accurately reflect the potential influence of the root cause node compared with GPS, especially when dealing with long-tail distribution data, avoiding the misjudgment problem caused by too large absolute difference.

[0166] When searching for the root cause dimension combination of a cluster from all the second candidate non-leaf node dimension combinations in the cluster according to the comprehensive evaluation index, it is possible to traverse and search according to the level of the dimension combination, and sort the second candidate non-leaf node dimension combinations. If it is found during the traversal of a certain level that the RS value of a second candidate non-leaf node dimension combination is greater than or equal to the threshold δ, then it may be the root cause dimension combination, stop the search for the current level and enter the next level for further analysis. Finally, the found root cause dimension combination satisfies two points: 1) expressiveness, that is, the root cause candidate set S can accurately represent the scope of the fault in the multi-dimensional data; 2) interpretability, that is, the root cause candidate set S is as concise as possible to facilitate the operator to focus on the relevant attributes and attribute values. In the case of multiple root causes coexisting, each root cause candidate set should have expressiveness and interpretability and not interfere with each other.

[0167] In addition, in order to ensure that the final root cause location result is concise and interpretable, the present invention applies the Occam's razor principle, only retaining the most interpretable and concise dimension combination and removing redundant information. For example, when two root causes (Guangdong, ChinaUnicom, *) and (*, ChinaUnicom, *) are recommended in the cluster, the system will retain the more concise (*, ChinaUnicom, *) to improve the interpretability and conciseness of the result.

[0168] In a specific example, the present invention shows an abnormal index root cause location algorithm for implementing the process of step S60. The specific algorithm process is as follows:

[0169]

[0170]

[0171] In summary, the method of the present invention is divided into two stages. In the performance metric anomaly detection stage, based on the AD metric of historical data and real-time deviation analysis, significant anomalies in the system are quickly identified, and their severity is quantified. Based on the AD metric, significant anomaly data is quickly screened out from a large amount of operation and maintenance data, narrowing the search scope of anomalies. The average time-consuming for this significant anomaly detection is only one second, ensuring high efficiency and real-time response ability. Subsequently, an autoencoder neural network is used to reconstruct the screened normal data to detect subtle but important anomalies, improving the detection ability for subtle and critical anomalies. Due to the efficient representation of normal patterns in its low-dimensional latent space and the amplification of the reconstruction error of abnormal data, the autoencoder can accurately capture subtle anomalies, avoiding the missed detection problem that may be caused by simply relying on the AD metric. In the anomaly metric root cause location stage, when dealing with the problem of uneven long-tail distribution of high-dimensional data, the present invention proposes an IPS scoring function based on CRE. By calculating the proportion of the difference between the actual value and the predicted value to the actual value, the possibility of an attribute combination as the root cause is evaluated, effectively reducing the impact of extreme values on the evaluation result when the long-tail distribution is uneven. In the calculation process of IPS, a certain weight balance is applied to low-frequency combinations, focusing on relative differences rather than absolute differences, which can more accurately reflect the actual influence of the root cause combination and reduce the interference of extreme values on the evaluation result. The proposed IPS scoring function. In addition, the present invention introduces a two-layer filtering mechanism of CF and LF. While maintaining the generalization ability of the model, it focuses on the fault area and preferentially evaluates high-risk data blocks, ensuring that while maintaining the robustness of the model, it focuses on high-risk areas. Compared with existing anomaly root cause location algorithms, the method of the present invention can not only more accurately measure the severity of anomalies, but also more effectively address the long-tail distribution problem in high-dimensional data, thereby improving the accuracy of root cause location.

[0172] To evaluate the effectiveness of the method of the present invention, experiments were conducted on two real-world datasets (a certain online shopping platform and a certain Internet company). The experimental results verified that the present invention achieved an improvement in the accuracy of root cause location compared to the baseline method. The experimental process and results are described below.

[0173] Experimental settings:

[0174] Regarding the dataset settings, to measure the effectiveness of the present invention (denoted as AL) in different scenarios, two real-world datasets were used in the experiment. L1 came from a certain online shopping platform, and L2 came from a certain Internet company. Table 2 summarizes the key statistical information of the datasets, where p represents the number of attributes. The datasets contain basic metrics (F) and derived metrics (D), both of which are gold standard signals. Among them, the average prediction residual is used to measure the deviation degree of normal leaf attribute combinations, and the explanatory power (EP) represents the proportion of the prediction residual of abnormal leaf attribute combinations to the overall residual, reflecting the relative significance of anomalies.

[0175] Table 2: Introduction to the dataset

[0176]

[0177] Regarding the evaluation metrics, the F1-score is used to evaluate the root cause localization of multi-dimensional data. The F1-Score is the harmonic mean of precision and recall, which is used to comprehensively evaluate the performance of the algorithm. When there is a trade-off between precision and recall, the F1-Score can balance the effects of precision and recall. Its calculation formula is:

[0178]

[0179] Among them, precision (also known as the positive predictive value) represents the proportion of samples predicted as positive classes that are actually positive classes. In this specification, it represents the proportion of samples predicted as faulty that actually have faults. The calculation formula is:

[0180]

[0181] Among them, TP is the number of samples correctly predicted as positive classes, and FP represents the number of negative class samples wrongly predicted as positive classes.

[0182] Recall (also known as the recall rate) represents the proportion of all positive examples that are correctly predicted, reflecting the detection ability of the algorithm in the true root cause dimension combination, that is, the sensitivity of the algorithm. Its calculation formula is:

[0183]

[0184] Among them, FN represents the number of positive class samples wrongly predicted as negative classes.

[0185] To comprehensively evaluate the performance of the present invention in multiple scenarios, multiple groups of experiments are designed based on existing research. The experiments are set for different root causes, including the varying number of elements n and the configuration of the cuboid layer, to conduct a detailed evaluation of the F1-Score.

[0186] The method of the present invention is compared with the following baseline methods:

[0187] ADtributor (ADT): Assume that the root cause is limited to a single dimension, and quantify the root cause through "explanatory power" and "surprise". Use the ARMA model to predict the KPI. First, rank the dimensions according to the surprise of the elements within the dimension, and then calculate the explanatory power of the elements. When the sum exceeds the threshold, these elements are regarded as the root cause.

[0188] Apriori (APR): Based on the principle that all subsets of frequent item sets must be frequent, the algorithm gradually increases the size of the item set and iteratively generates frequent item sets that meet the minimum support threshold. By using a pruning strategy to eliminate item sets that are unlikely to be frequent, it effectively reduces the number of candidate item sets, calculates the support for the remaining item sets, and only retains the item sets that meet the frequency condition until no new frequent item sets can be generated. Lin et al. applied this method to the root cause localization of abnormal leaf attribute combinations, using the Apriori algorithm combined with confidence evaluation to mine association rules, thereby effectively identifying potential root causes.

[0189] HotSpot (HS): This method proposes a rippling effect root cause determination method based on prediction and search strategies, and uses Potential Score to quantify the relationship between dimension combinations and leaf nodes. Different from ADtributor, HotSpot considers multi-dimensional root cause attribute combinations and uses Monte Carlo Tree Search (MCTS) and pruning strategies to find the set of attribute combinations with the highest potential score.

[0190] ImpAPTr (IAP): It uses breadth-first search to locate the attribute combinations that maximize the impact factor and diversity factor. Since ImpAPTr only ranks the attribute combinations instead of directly determining the root cause attribute combinations, the top n attribute combinations are selected as the root cause combinations. In addition, considering that the original impact factor has an inhibitory effect on the measured value, an adaptive method is used to dynamically adjust the impact factor for each fault.

[0191] MID: It locates the root cause by searching for the attribute combination that maximizes the objective function and uses an entropy-based heuristic method to accelerate the search. Since its objective function is limited to specific scenarios and performs poorly on general multi-dimensional data, its objective function is replaced by IPS.

[0192] Squeeze (SQ): Squeeze proposes GRE based on HotSpot, extends it to derived metrics such as success rate, and solves the problem of 0 prediction values. Different from previous algorithms, Squeeze first prunes normal leaf nodes from bottom to top to reduce the search space, and then clusters the leaf nodes according to the deviation score. Subsequently, it searches for the root cause within each cluster from top to bottom and quantifies the root cause through GPS.

[0193] PSqueeze (PSQ): The PSqueeze (PSQ) method performs probabilistic clustering based on GRE, simplifies the multi-dimensional root cause problem to a single root cause, and applies the GPS heuristic method within each cluster for efficient search. Finally, PSqueeze determines whether there is an external root cause by evaluating the GPS score for root cause localization.

[0194] The moving average (MA) method is used for prediction. The predicted value of leaf attribute combination e at a specific time t0 is estimated by calculating the average value of e from time t-10 to t-1. The reason for choosing MA is its simplicity and low computational cost. Table 3 shows the time consumed by several algorithms for a single leaf attribute combination.

[0195] Table 3 Comparison of the usage time of prediction methods

[0196]

[0197] As Figure 5 shown, under the conditions of different numbers of elements n and cuboid level settings, AL outperforms the baseline model in terms of performance in six different datasets and achieves the best results. Tables 4 and 5 further show that although the overall performance of AL slightly decreases when the prediction residuals increase (datasets B5 to B8), its performance is still better than that of traditional methods. The results show that this method has strong robustness in multi-dimensional root cause identification of faults. The performance of ADT is limited by the increase in the number of cuboid layers. Especially in deep cuboids, its performance will significantly decrease due to smaller abnormal magnitudes and increased complexity of attribute combinations. Although APR shows excellent performance in multiple experimental settings and reaches the best results in certain specific settings, its high sensitivity to parameters leads to very unstable performance in some cases. For example, in dataset A and datasets B5 - B8, when the number of cuboid layers is 3, the performance of APR significantly decreases. This volatility stems from the fact that the hierarchical pruning strategy of APR may wrongly prune the correct search path. Especially when the number of cuboid layers is large, the error of the pruning strategy will intensify, resulting in information loss and limiting the effectiveness of the algorithm. Although HS is less affected by the change of prediction residuals, its performance is poor. Both MID and IAP have limitations for specific metrics. Especially when considering multiple root causes, complex data structures, and small abnormal magnitudes, their performance is significantly limited. The heuristic search strategy of MID is highly dependent on the scenario and it is difficult to adapt to deeper cuboid structures. While the sensitivity of IAP to attribute combinations makes it perform poorly when dealing with small abnormal magnitudes. In the root cause localization task, compared with SQ and PSQ, AL shows stronger robustness and is better than the baseline method. The F1 score of AL in all test scenarios is increased by about 6% on average, exceeding PSQ. The core reason for this performance improvement is that SQ and PSQ are difficult to accurately capture tiny anomalies and are easily affected by the uneven long-tail distribution of multi-dimensional data, resulting in low localization efficiency.

[0198] Elaborate on the superiority of AL (the present invention) in detail from two aspects: the adaptability of the long-tail distribution in the anomaly detection stage and the root cause localization stage, and at the same time analyze its performance differences on different datasets. First, in the anomaly detection stage, AL introduces a two-stage detection framework of the AD index and the autoencoder to replace the single prediction residual judgment mechanism in traditional methods. Traditional methods only judge anomalies by the size of the prediction residual, which is prone to missed detections and cannot effectively detect slight but important anomaly signals. The two-stage detection framework of AL first uses the AD index to quickly screen significant anomalies in large-scale data with a low anomaly rate. The AD index can more finely evaluate the anomaly nature of data points. Especially for those anomalies with a large prediction residual, it can better distinguish different degrees of anomalies. After the preliminary screening, the autoencoder model will further reconstruct the normal data to capture subtle but critical anomalies. In contrast, PSQ uses a single clustering method, which is easily affected by noise and large prediction residuals in multi-dimensional data, resulting in a decline in anomaly detection performance. This two-stage strategy not only improves the sensitivity of anomaly detection but also significantly reduces false alarms, ensuring the detection accuracy.

[0199] Secondly, in the root cause localization stage, AL introduces the IPS combined with the collaborative filtering mechanism to address the imbalance problem of the long-tail distribution in multi-dimensional data. The traditional CPS method is easily interfered by low-frequency extreme values on datasets with uneven long-tail distributions. This is because in multi-dimensional data, low-frequency extreme values usually correspond to special attribute combinations. When the frequencies of these combinations are low but the anomaly degrees are high, the CPS method is prone to false alarms under the influence of extreme values. In contrast, IPS applies a certain weight balance to low-frequency combinations during the calculation process, identifies the root causes of abnormal fluctuations in key features or dimensions, more accurately reflects the actual influence of the root cause combination, successfully avoids the impact of the long-tail distribution on performance, and reduces the interference of extreme values on the evaluation results.

[0200] In datasets A and D, the complexity of the attribute combination is low, and the long-tail distribution characteristics are relatively weak. Compared with the baseline method, the two-stage detection framework of AL enables AL to effectively detect and locate significant anomalies and subtle anomalies. As shown in Table 5, AL and many other benchmarks rely on predicted values. In datasets with large prediction residuals (B7 and B8) or in cases with a large number of elements and a high number of cuboid layers, the F1-Score of AL decreases, but it is still better than other methods. The attribute combinations of these datasets are more complex, the long-tail distribution characteristics are more significant, and the interference of extreme values is large. The IPS scoring method can effectively balance the influence of low-frequency extreme values, making AL still perform well on these datasets, indicating the applicability and robustness of the IPS scoring in long-tail distribution data and its suitability for complex multi-factor data environments.

[0201] Table 4 Overall performance in datasets A and D

[0202]

[0203]

[0204] Table 5: Comparison of Overall Performance in Datasets B5, B6, B7 and B8

[0205]

[0206]

[0207] Influence of AD, Autoencoder and IPS:

[0208] To verify the effectiveness of each sub-module in AL, six groups of ablation experiments were designed to clarify the contribution of each module to the overall performance. As Figures 6 to 11 shown, the performance of all variants is lower than that of the AL model to varying degrees, but better than the baseline method, indicating that each sub-module of AL can effectively improve the root cause localization ability for multi-dimensional data. Among them, Figure 6 is the F1-score of each algorithm for root cause localization only based on the degree of anomaly, Figure 7 is the F1-score of each algorithm for root cause localization only based on the data reconstruction error, Figure 8 is the F1-score of each algorithm for root cause localization only based on IPS, Figure 9 is the F1-score of each algorithm for root cause localization based on the degree of anomaly and the data reconstruction error, Figure 10 is the F1-score of each algorithm for root cause localization based on the degree of anomaly and IPS, Figure 11 is the F1-score of each algorithm for root cause localization based on the data reconstruction error and IPS.

[0209] In the ablation experiment analysis of AL, the contributions of different modules to the overall performance of AL were evaluated in detail. It was found that using the AD module alone could quickly identify significant anomalies, but its ability to detect minor anomalies was relatively weak, so the overall performance was limited. This indicates that although the AD module is efficient, it is not sufficient to handle all anomaly situations independently. When only using the autoencoder neural network, by reconstructing normal data, the autoencoder can learn the normal patterns of the data, become highly sensitive to subtle changes, and improve the ability to detect minor anomalies. The IPS mechanism improves the accuracy of anomaly localization by reducing the influence of extreme values in high-dimensional data. In the combined experiment of AD and the autoencoder, the two modules complement each other and exhibit good complementarity. The AD module can quickly lock the significant anomaly regions, while the autoencoder deeply excavates and detects minor anomalies within these regions, thus achieving comprehensive and accurate anomaly detection. The combination of AD and IPS performs excellently when the long-tailed distribution of high-dimensional data is uneven, outperforming the baseline method. The joint use of the autoencoder and IPS further verifies their advantages in identifying minor anomalies and dealing with uneven long-tailed data distributions. The synergistic effect of each module improves the system performance.

[0210] Through the ablation experiment analysis, it can be confirmed that each module of AL contributes to the improvement of the overall performance, and there is complementarity between the modules, effectively enhancing the root cause localization ability of AL at different levels. Compared with the baseline method, the present invention can comprehensively detect anomalies of different degrees and effectively cope with the interference of extreme values when the long-tailed distribution of high-dimensional data is uneven. The F1 score is significantly better than the baseline. This result verifies the rationality of the model architecture proposed by the present invention and the practical application potential of this method in root cause localization in complex microservice architectures.

[0211] Regarding the computational efficiency of AL, all experiments were conducted on a server equipped with an NVIDIA GeForce RTX 3060 graphics card, with an Intel Xeon E5-2620 v3 processor of x86_64 architecture, configured with 32 physical cores, and the operating system was Linux 5.4.0-155-generic. The experimental environment used CUDA 11.3 and driver version 535.86.10. All algorithms were implemented in Python and relied on mature open-source libraries such as Pandas and NumPy. The experimental results are shown in Figures 12 to 17 . Among them, Figure 12 is the comparison result of the single-fault running time of each algorithm on the A dataset, Figure 13 is the comparison result of the single-fault running time of each algorithm on the D dataset, Figure 14 is the comparison result of the single-fault running time of each algorithm on the B5 dataset, Figure 15 is the comparison result of the single-fault running time of each algorithm on the B6 dataset, Figure 16The comparison results of the single - fault running time of each algorithm on the B7 dataset Figure 17 The comparison results of the single - fault running time of each algorithm on the B8 dataset. It can be seen from the experimental results that under the same experimental conditions, the method AL of the present invention shows obvious advantages in terms of running time. Even in the worst case, the time overhead of single - fault detection remains within 10 seconds, meeting the real - time requirements of the actual system. The high efficiency and accuracy of AL benefit from the fine - grained anomaly detection mechanism ensuring high detection accuracy, the heuristic search strategy effectively coping with long - tail distribution data, and the collaborative filtering mechanism significantly reducing the time overhead, enabling it to quickly locate the root cause and be applicable to complex scenarios. In contrast, the baseline methods all have deficiencies in performance. Although ADT has an extremely fast running time on datasets A, B5 - B8, its single - attribute assumption leads to low detection accuracy, and its running time on dataset D reaches hundreds of seconds, making it difficult to adapt to the derived - metric scenario. IAP uses a breadth - first search strategy. Although its running time is close to that of AL in some cases, its attribute - combination sorting mechanism increases the operation complexity, resulting in additional time overhead. PSQ and SQ rely on bottom - up clustering and top - down localization processes, which cause more time and resource consumption in high - dimensional long - tail distribution data, and the average anomaly - localization time is higher than that of AL. MID usually takes dozens of seconds, but its detection accuracy is poor, mainly applicable to specific scenarios and lacking generality. The time overhead of HS increases exponentially with the increase of data complexity and cannot meet the requirements of real - time detection. APR is based on the association - rule mining algorithm, and the complexity of rule mining makes it perform particularly poorly in high - dimensional data scenarios, with a running time reaching hundreds of seconds and being difficult to meet the real - time requirements of the actual system. Generally speaking, the running time and detection accuracy of AL are better than those of other methods on each dataset.

[0212] Regarding the performance of AL under different configurations, except for the IPS threshold δ, the remaining parameters are automatically configured. Figure 18It shows the F1-score of the present invention under different δ values when the number of cuboid layers = 3 and n_element = 3. This setting is chosen because it is the most difficult one. In the experiment, the changes in the F1 scores of datasets B5, B6, B7, and B8 were compared as the threshold gradually increased from 0.25 to 0.9. From the data in the figure, it can be observed that as the threshold increases, the performance of the model on each dataset steadily improves and tends to be stable and reach a high value at a threshold of 0.9. Especially in the range where the threshold is from 0.85 to 0.9, the changes in various indicators tend to be small, indicating that the performance of the model is close to the optimal state. In root cause localization, a higher threshold usually means stricter filtering conditions, which helps to reduce the missed detection rate. Therefore, the model under this threshold can effectively locate the root cause of anomalies. In addition, the IPS threshold of 0.9 shows high F1-score stability on different datasets, indicating that this threshold has good adaptability and generalization ability and can achieve consistent performance in multiple scenarios. Therefore, setting the IPS threshold to 0.9 not only optimizes the accuracy and robustness of the model but also takes into account the actual application requirements, which is a reasonable and preferred setting. The experimental results in Table 4 and Table 5 also show that δ = 0.9 can achieve higher efficiency.

[0213] The experimental results show that on a real dataset containing 5400 faults, the F1 score of the present invention is increased by 6% compared with the existing method, improving the accuracy and robustness of root cause localization. The experimental results verify the advantages of the present invention compared with the existing methods.

[0214] In summary, in response to the root cause localization challenges brought by the high dynamism and complexity of the microservice architecture, the present invention proposes a robust root cause localization method, which overcomes the deficiencies of existing methods in detecting different degrees of anomalies and dealing with the long-tail distribution of multi-dimensional data. The present invention quickly screens significant anomalies in large-scale low-anomaly-rate data by introducing the AD index and combines an autoencoder neural network to reconstruct normal data, thereby effectively capturing subtle and critical anomalies and reducing the risk of missed detection. In addition, the IPS mechanism combined with a two-layer filtering strategy shows stronger robustness when the long-tail distribution of high-dimensional data is unbalanced, accurately focuses on the fault area, and effectively reduces the interference of extreme values on fault localization. The experimental results show that the F1 score of the present invention is increased by 6% compared with the existing optimal method in multiple application scenarios, and it is more robust especially in scenarios where the influence of extreme values is significant.

[0215] The method provided by the embodiment of the present invention can be applied to an electronic device. Specifically, the electronic device can be: a desktop computer, a portable computer, a smart mobile terminal, a server, etc. There is no limitation here. Any electronic device that can implement the present invention belongs to the protection scope of the present invention.

[0216] Based on the same inventive concept, an embodiment of the present invention further provides a multi-dimensional data root cause localization device based on reconstruction and heuristic search, including:

[0217] An acquisition module, configured to acquire the observed values of multi-dimensional data, and perform data prediction according to the observed values to obtain the predicted values of the multi-dimensional data;

[0218] A calculation module, configured to calculate the abnormality degree value of each leaf node dimension combination according to the observed values and the predicted values; the leaf node dimension combination is the dimension combination corresponding to the finest-grained data in the multi-dimensional data;

[0219] A first screening module, configured to screen out significantly abnormal data and non-significantly abnormal data from the multi-dimensional data according to the abnormality degree value;

[0220] A reconstruction module, configured to use a pre-trained auto-encoder neural network to reconstruct the non-significantly abnormal data and calculate the reconstruction error;

[0221] A second screening module, configured to further screen out subtle key abnormal data from the non-significantly abnormal data according to the reconstruction error; the abnormality of the subtle key abnormal data is more subtle than that of the significantly abnormal data and belongs to key abnormality;

[0222] A localization module, configured to perform multi-dimensional data root cause localization by using a heuristic search mechanism according to the significantly abnormal data and the subtle key abnormal data.

[0223] An embodiment of the present invention further provides an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus.

[0224] The memory is used to store a computer program;

[0225] The processor, when executing the program stored in the memory, implements the method steps of any one of the above multi-dimensional data root cause localization based on reconstruction and heuristic search.

[0226] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only a thick line is used in the figure to represent it, but it does not mean that there is only one bus or one type of bus.

[0227] The communication interface is used for communication between the above-mentioned electronic device and other devices.

[0228] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0229] The aforementioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0230] It should be noted that for the apparatus / electronic device embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, please refer to the partial description of the method embodiments.

[0231] It should be noted that the terms "first", "second", etc. are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention.

[0232] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.

[0233] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A multi-dimensional data root cause location method based on reconstruction and heuristic search, characterized in that Including: Obtain the observations of multi-dimensional data, and perform data prediction based on the observations to obtain the predicted values of the multi-dimensional data; Calculate the anomaly degree value of each leaf node dimension combination according to the observations and the predicted values; the leaf node dimension combination is the dimension combination corresponding to the finest-grained data in the multi-dimensional data; Screen significant abnormal data and non-significant abnormal data from the multi-dimensional data according to the anomaly degree value; Use a pre-trained autoencoder neural network to reconstruct the non-significant abnormal data and calculate the reconstruction error; Continue to screen subtle critical abnormal data from the non-significant abnormal data according to the reconstruction error; the anomaly of the subtle critical abnormal data is more subtle than that of the significant abnormal data and belongs to critical anomalies; Perform root cause localization of multi-dimensional data using a heuristic search mechanism according to the significant abnormal data and the subtle critical abnormal data.

2. The multi-dimensional data root cause location method based on reconstruction and heuristic search according to claim 1, wherein Calculating the anomaly degree value of each leaf node dimension combination according to the observations and the predicted values includes: Among them, v d (ε) represents the observed value of the leaf node dimension combination ε for which the abnormal degree value is currently to be calculated, and f d (ε) represents the predicted value of the leaf node dimension combination ε; v d (ε i ) represents the observed values of other leaf node dimension combinations ε i at the same moment, and f d (ε i ) represents the predicted values of other leaf node dimension combinations ε i . n is the number of other leaf node dimension combinations at the same moment, and AD(ε) is the abnormal degree value of the leaf node dimension combination ε.

3. The multi-dimensional data root cause location method based on reconstruction and heuristic search according to claim 1, characterized in that, The observations of the multi-dimensional data include observations at multiple time points; The data prediction according to the observations to obtain the predicted values of the multi-dimensional data includes: Use a sliding time window to segment the observations at the multiple time points; Take the average of the observations at the time points within each sliding time window to obtain the predicted values at each time point within the sliding time window.

4. The multi-dimensional data root cause localization method based on reconstruction and heuristic search according to claim 1, characterized in that The screening of significant abnormal data and non-significant abnormal data from the multi-dimensional data according to the anomaly degree value includes: Obtain the CDF curve of the anomaly degree value; Determine the first screening threshold using the knee point method according to the CDF curve; Use the first screening threshold to screen significant abnormal data and non-significant abnormal data from the multi-dimensional data.

5. The multi-dimensional data root cause location method based on reconstruction and heuristic search according to claim 1, characterized in that The pre-trained autoencoder neural network includes an encoder and a decoder; The use of the pre-trained autoencoder neural network to reconstruct the non-significant abnormal data is specifically to input the non-significant abnormal data into the pre-trained autoencoder neural network to achieve the following operations: The encoder compresses the non-significant abnormal data through a linear transformation and extracts the low-dimensional latent feature representation of the non-significant abnormal data: z = f θ (x) = σ(W e x + b e ); where z represents the low-dimensional latent feature representation, denotes the non-significant abnormal data of dimension d in the real number field , f θ (·) represents the encoder, is the weight matrix of the encoder, k << d, is the bias vector of the encoder; The decoder reconstructs the non-significant abnormal data according to the low-dimensional latent feature representation: Among them, represents the reconstructed data, g θ' (·) represents the decoder, σ(·) is the nonlinear activation function Sigmoid, is the weight matrix of the decoder, is the bias vector of the decoder.

6. The multi-dimensional data root cause location method based on reconstruction and heuristic search according to claim 5, characterized in that The continued screening of subtle critical abnormal data from the non-significant abnormal data according to the reconstruction error includes: Calculate the mean absolute error MAE of the reconstruction error; Use the kernel density estimation method to model the distribution of the mean absolute error MAE to obtain the MAE distribution; Determine the second screening threshold using the knee point method according to the MAE distribution; Use the second screening threshold to continue to screen subtle critical abnormal data from the non-significant abnormal data.

7. The multi-dimensional data root cause location method based on reconstruction and heuristic search according to claim 1, characterized in that The root cause localization of multi-dimensional data using a heuristic search mechanism according to the significant abnormal data and the subtle critical abnormal data includes: Merge the significant abnormal data and the subtle critical abnormal data into an abnormal data set; Cluster the abnormal data set through a generalized ripple effect to obtain multiple clusters; For each cluster, calculate the decline scores of all non-leaf node dimension combinations within the cluster, and filter multiple first candidate non-leaf node dimension combinations based on the decline scores; calculate the internal influence degree and the cluster influence degree of each first candidate non-leaf node dimension combination; if the internal influence degree and the cluster influence degree of a first candidate non-leaf node dimension combination are both higher than the corresponding thresholds, then use it as a second candidate non-leaf node dimension combination, and calculate its potential score; generate a comprehensive evaluation index based on the internal influence degree, the cluster influence degree and the potential score of each second candidate non-leaf node dimension combination; according to the comprehensive evaluation index, find the root cause dimension combination of the cluster from all the second candidate non-leaf node dimension combinations of the cluster; wherein, the potential score is used to evaluate the abnormality degree of the non-leaf node dimension combination by using relative difference.

8. The multi-dimensional data root cause location method based on reconstruction and heuristic search according to claim 7, characterized in that The calculation method of the potential score is as follows: Wherein, v(e1) represents the observed value of the leaf node dimension combination under the second candidate non-leaf node dimension combination for which the potential score is to be calculated, a(e1) represents the expected value of the leaf node dimension combination under the second candidate non-leaf node dimension combination, f(e1) represents the observed values of other leaf node dimension combinations in the cluster where the second candidate non-leaf node dimension combination is located, f(e2) represents the predicted values of other leaf node dimension combinations in the cluster where the second candidate non-leaf node dimension combination is located, avg(·) represents taking the average of the calculation results in (·) among the leaf node dimension combinations, and IPS represents the potential score.

9. A multi-dimensional data root cause location device based on reconstruction and heuristic search, characterized in that Including: An acquisition module, configured to acquire the observed values of multi-dimensional data, and perform data prediction according to the observed values to obtain the predicted values of the multi-dimensional data; A calculation module, configured to calculate the abnormality degree value of each leaf node dimension combination according to the observed values and the predicted values; the leaf node dimension combination is the dimension combination corresponding to the finest-grained data in the multi-dimensional data; A first screening module, configured to screen out significantly abnormal data and non-significantly abnormal data from the multi-dimensional data according to the abnormality degree value; A reconstruction module, configured to perform data reconstruction on the non-significantly abnormal data by using a pre-trained autoencoder neural network, and calculate the reconstruction error; A second screening module, configured to further screen out subtle key abnormal data from the non-significantly abnormal data according to the reconstruction error; the abnormality of the subtle key abnormal data is more subtle than that of the significantly abnormal data and belongs to key abnormality; A positioning module, configured to perform root cause localization of multi-dimensional data by using a heuristic search mechanism according to the significantly abnormal data and the subtle key abnormal data.

10. An electronic device, characterized in that, Including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory complete communication with each other through the communication bus, The memory is used for storing a computer program; The processor is configured to implement the multi-dimensional data root cause localization method based on reconstruction and heuristic search according to any one of claims 1 to 8 when executing the computer program stored on the memory.