Fault root cause positioning method in micro-service system and related device
By detecting the abnormality of the key business indicators of microservice nodes in the microservice system and building a fault impact diagram for reverse tracking, the root cause location problem in the microservice system is solved, and fast and accurate fault location is achieved.
Patent Information
- Application Number
- CN202510519374.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-05-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In microservice systems, as the number of microservice nodes increases and the complexity of interaction increases, failure of a single microservice node may cause cascading failures, resulting in system performance degradation or service interruption. How to quickly and accurately locate the root cause of failures has become an urgent problem to be solved.
By performing performance abnormality detection on the key business indicators of each microservice node in the microservice system, when an outlier is detected, a fault impact graph that simulates the fault propagation path between microservice nodes is constructed, and the graph is traversed along the backpropagation path, evaluating the suspicion score of each microservice node, and finally determining the target microservice node with the highest suspicion score is the root cause node of the failure.
It realizes the rapid and accurate positioning of the root cause node when a fault occurs. Through intuitive fault impact graph and reverse tracking technology, the suspicion score of each microservice node is comprehensively evaluated, so as to quickly filter out the final root cause node.
Smart Images

Figure CN120034425A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of fault location technology, and more specifically, to a method for locating a root cause of a fault in a microservice system and a related device. Background Art
[0002] With the continuous development of cloud computing and cloud native technology, microservice systems are becoming more and more popular due to their flexibility, scalability and agility. Microservice applications decompose services into a set of independent microservice nodes, allowing each microservice node to be independently developed and updated, and interact and share data through the network to collaboratively provide complete services.
[0003] However, as the number of microservice nodes increases and the complexity of interactions increases, a single microservice node failure may spread, leading to cascading failures, which in turn causes overall system performance degradation or service interruption. Therefore, how to quickly and accurately locate the root cause of a failure when it occurs has become a technical problem that technicians in this field need to solve urgently. Summary of the invention
[0004] In view of this, the present invention discloses a method for locating the root cause of a fault in a microservice system and a related device, so as to quickly and accurately locate the root cause of the fault when a fault occurs.
[0005] A method for locating the root cause of a fault in a microservice system, comprising:
[0006] Perform performance anomaly detection on key business indicators of each microservice node in the microservice system;
[0007] When an abnormal value is detected from all the key business indicators, a fault impact graph simulating a fault propagation path between microservice nodes is constructed based on the abnormal value;
[0008] The microservice node corresponding to the abnormal value is taken as an abnormal node, and the fault impact graph is traversed along the reverse propagation path starting from the abnormal node to evaluate the suspicion score of each microservice node;
[0009] The target microservice node with the highest suspicion score is identified as the fault root cause node.
[0010] Optionally, the performing performance anomaly detection on key business indicators of each microservice node in the microservice system includes:
[0011] The isolation forest algorithm is used to perform performance anomaly detection on the key business indicators of each microservice node in the microservice system.
[0012] Optionally, when an abnormal value is detected from all the key business indicators, a fault impact graph simulating a fault propagation path between microservice nodes is constructed based on the abnormal value, including:
[0013] When an abnormal value is detected from all the key business indicators, a fault propagation path is extracted from the call relationship between the microservice nodes within the fault propagation time window;
[0014] Constructing a fault impact graph skeleton based on each of the fault propagation paths;
[0015] Determining a fault propagation probability of each of the fault propagation paths in the fault impact graph skeleton;
[0016] The fault propagation probability is added to the corresponding fault propagation path on the fault impact diagram skeleton to obtain the fault impact diagram.
[0017] Optionally, determining the fault propagation probability of each fault propagation path in the fault impact graph skeleton includes:
[0018] extracting a fault propagation feature of each of the fault propagation paths;
[0019] The fault propagation characteristics are processed using principal component analysis technology to obtain the fault propagation probability corresponding to the fault propagation path.
[0020] Optionally, the fault propagation characteristics include: node abnormality, node activity, intimacy and similarity between two microservice nodes in the fault propagation path.
[0021] Optionally, taking the microservice node corresponding to the abnormal value as an abnormal node, traversing the fault impact graph along a reverse propagation path starting from the abnormal node, and evaluating the suspicion score of each microservice node includes:
[0022] Reversing the edge representing the fault propagation path in the fault impact graph to obtain a reversed fault impact graph;
[0023] The reverse fault impact graph is traversed along a reverse propagation path starting from the abnormal node, and the suspicion score of each of the microservice nodes is evaluated.
[0024] Optionally, traversing the inverse fault impact graph along the reverse propagation path starting from the abnormal node to evaluate the suspicion score of each microservice node includes:
[0025] Based on the exhaustive traversal method, the inverse fault impact graph is traversed along the reverse propagation path starting from the abnormal node, and the inverse process of fault propagation is simulated by the suspicion score iteration method to evaluate and obtain the suspicion score of each microservice node.
[0026] A device for locating root cause of a fault in a microservice system, comprising:
[0027] The anomaly detection unit is used to detect performance anomalies of key business indicators of each microservice node in the microservice system;
[0028] A fault impact graph construction unit is used to construct a fault impact graph simulating a fault propagation path between microservice nodes based on the abnormal value when an abnormal value is detected from all the key business indicators;
[0029] A score evaluation unit, configured to take the microservice node corresponding to the abnormal value as an abnormal node, traverse the fault impact graph along a reverse propagation path starting from the abnormal node, and evaluate the suspicion score of each microservice node;
[0030] The root cause node determination unit is used to determine the target microservice node with the highest suspicion score as the fault root cause node.
[0031] A computer storage medium stores at least one instruction, and when the at least one instruction is executed by a processor, any of the above-mentioned methods for locating the root cause of a fault in a microservice system is implemented.
[0032] An electronic device, comprising: a memory and a processor;
[0033] The memory is used to store at least one instruction;
[0034] The processor is used to execute the at least one instruction to implement any of the above-mentioned methods for locating the root cause of a fault in a microservice system.
[0035] As can be seen from the above technical scheme, the present invention discloses a method and related device for locating the root cause of a fault in a microservice system, and performs performance anomaly detection on the key business indicators of each microservice node in the microservice system. When an outlier is detected from all the key business indicators, a fault impact graph simulating the fault propagation path between microservice nodes is constructed based on the outlier, and the microservice node corresponding to the outlier is taken as an abnormal node. The fault impact graph is traversed along the reverse propagation path from the abnormal node to evaluate the suspicion score of each microservice node, and the target microservice node with the highest suspicion score is determined as the root cause node of the fault. The present application simulates the fault propagation path between microservice nodes by constructing a fault impact graph, so as to intuitively present all potential fault propagation paths; by adopting reverse tracing technology, the entire fault impact graph is traversed along the reverse propagation path from the abnormal node to comprehensively evaluate the suspicion scores of each microservice node, so as to quickly screen out the microservice node with the highest suspicion score from all suspicion scores as the final root cause node of the fault. Therefore, the present application realizes the rapid and accurate positioning of the root cause node of the fault when a fault occurs. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the disclosed drawings without paying creative work.
[0037] Figure 1 A flow chart of a method for locating the root cause of a fault in a microservice system disclosed in an embodiment of the present invention;
[0038] Figure 2 A flow chart of a method for constructing a fault impact graph disclosed in an embodiment of the present invention;
[0039] Figure 3 A schematic diagram of a fault impact graph skeleton disclosed in an embodiment of the present invention;
[0040] Figure 4 A technical flow chart of locating the root cause of a fault in a microservice system disclosed in an embodiment of the present invention;
[0041] Figure 5 A schematic diagram of the structure of a device for locating the root cause of a fault in a microservice system disclosed in an embodiment of the present invention;
[0042] Figure 6 The present invention is a schematic diagram of the structure of an electronic device disclosed in an embodiment of the present invention. DETAILED DESCRIPTION
[0043] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0044] The embodiment of the present invention discloses a method and related device for locating the root cause of a fault in a microservice system. The method simulates the fault propagation path between microservice nodes by constructing a fault impact graph, so as to intuitively present all potential fault propagation paths. The reverse tracing technology is adopted to traverse the entire fault impact graph along the reverse propagation path starting from the abnormal node to comprehensively evaluate the suspicion score of each microservice node, so as to quickly screen out the microservice node with the highest suspicion score from all suspicion scores as the final root cause node of the fault. Therefore, the present application realizes the rapid and accurate locating of the root cause node of the fault when a fault occurs.
[0045] See also Figure 1 , a flow chart of a method for locating the root cause of a fault in a microservice system disclosed in an embodiment of the present application, the method may include:
[0046] Step S101: Perform performance anomaly detection on key business indicators of each microservice node in the microservice system.
[0047] In a highly dynamic microservice environment, even if no anomalies occur, service performance will have peaks and valleys, so it is difficult to distinguish between true anomalies and normal fluctuations. Based on this, this application performs performance anomaly detection on the key business indicators of each microservice node in the microservice system to promptly detect abnormal values in key business indicators.
[0048] Preferably, the present application uses the isolation forest algorithm to detect performance anomalies for the key business indicators of each microservice node in the microservice system, thereby achieving high accuracy of anomaly detection. The isolation forest algorithm is an anomaly detection algorithm based on a decision tree, which detects outliers in a data set by constructing a set of randomly partitioned binary trees. The basic idea is to regard abnormal samples as isolated points in the feature space, while normal samples are densely distributed in the feature space.
[0049] It should be noted that detecting the key business indicators of microservice nodes is an effective means to evaluate the health status of microservice nodes and discover problems. The key business indicators of each microservice node may include one or more, see the example shown in Table 1.
[0050] Table 1
[0051]
[0052] To prevent any potential failures from being missed, it is necessary to monitor the performance of all microservice nodes and use the isolation forest algorithm to detect performance anomalies for the key business indicators of each microservice node.
[0053] When a series of abnormal values are continuously detected in any microservice node, it can be determined that a performance abnormality has occurred in the microservice node, which is usually caused by a failure somewhere in the microservice node.
[0054] In order to measure the degree of deviation of each outlier a from the normal level, formula (1) can be used to calculate the anomaly score corresponding to the outlier a, so as to further analyze and locate the cause of the problem.
[0055] (1);
[0056] In the formula, represents the anomaly score corresponding to the outlier value a, The normalized value of the key business indicator representing the outlier a, Represents the normalized value of the key business indicator of the microservice node at timestamp t, represents the value of the key business indicator at timestamp t, T represents the time period before the abnormal value a occurs, Indicates the duration contained in time period T.
[0057] Step S102: When an abnormal value is detected from all the key business indicators, a fault impact graph simulating a fault propagation path between microservice nodes is constructed based on the abnormal value.
[0058] After detecting anomalies from key business indicators, this application will build a fault impact graph that simulates the fault propagation path between microservice nodes. The fault impact graph shows all possible fault propagation paths.
[0059] Step S103: taking the microservice node corresponding to the abnormal value as the abnormal node, traversing the fault impact graph along the reverse propagation path starting from the abnormal node, and evaluating the suspicion score of each microservice node.
[0060] In the fault impact graph, each node represents a microservice node, and the edge represents the fault propagation path between microservice nodes. Through the fault impact graph, we can directly understand the impact of a microservice node failure on other microservice nodes in the microservice system. This application takes the microservice node corresponding to the outlier as the abnormal node. Starting from the abnormal node, the process of tracing the fault source or the impact range along the reverse direction of the fault propagation path can be traced to determine the root cause node of the fault.
[0061] The suspicion score of a microservice node is a probability indicator used to measure the probability that a microservice node may have problems or failures in the microservice system. Starting from the abnormal node, while traversing the fault impact graph along the back propagation path, the suspicion score of each microservice node is evaluated to facilitate the subsequent determination of the final root cause node of the fault.
[0062] Step S104: determine the target microservice node with the highest suspicion score as the fault root cause node.
[0063] The highest suspicion score indicates that the corresponding microservice node has a higher probability index of a problem or failure. Therefore, the present application determines the target microservice node with the highest suspicion score as the root cause node of the failure.
[0064] In summary, the present application discloses a method for locating the root cause of a fault in a microservice system, and performs performance anomaly detection on the key business indicators of each microservice node in the microservice system. When an outlier is detected from all the key business indicators, a fault impact graph simulating the fault propagation path between microservice nodes is constructed based on the outlier, and the microservice node corresponding to the outlier is taken as an abnormal node. The fault impact graph is traversed along the reverse propagation path from the abnormal node, and the suspicion score of each microservice node is evaluated, and the target microservice node with the highest suspicion score is determined as the root cause node of the fault. The present application simulates the fault propagation path between microservice nodes by constructing a fault impact graph, and realizes the intuitive presentation of all potential fault propagation paths; by adopting the reverse tracing technology, the entire fault impact graph is traversed along the reverse propagation path from the abnormal node to comprehensively evaluate the suspicion scores of each microservice node, so as to quickly screen out the microservice node with the highest suspicion score from all suspicion scores as the final root cause node of the fault. Therefore, the present application realizes the rapid and accurate positioning of the root cause node of the fault when a fault occurs.
[0065] In one embodiment, see Figure 2 , a flowchart of a method for constructing a fault impact graph disclosed in an embodiment of the present application, that is, step S102 may specifically include:
[0066] Step S201: When an abnormal value is detected from all key business indicators, a fault propagation path is extracted from the call relationship between microservice nodes within the fault propagation time window.
[0067] Usually, in order to process the request, the upstream microservice node a will call the downstream microservice node b. After receiving the call request, the downstream microservice node b returns the corresponding result. During this period, the upstream microservice node a and the downstream microservice node b will affect each other. On the one hand, the failure of the upstream microservice node a (such as network delay) will lead to a decrease in the number of requests to the downstream microservice node b, resulting in abnormal transaction times. On the other hand, the CPU (Central Processing Unit) or memory failure of the downstream microservice node b will extend the response time of the upstream microservice node a and the downstream microservice node b, resulting in a decrease in the success rate of microservice node calls, which means that the failure can spread from the upstream to the downstream, and vice versa.
[0068] In order to capture all possible fault propagation paths, this application collects and analyzes the call relationships of microservice nodes within a period of time before the outlier is detected. During this period, the fault occurs and spreads to the entire microservice system. This application defines this period of time as the fault propagation time window (fpw) and adjusts it according to demand. For example, when a short-term sudden fault needs to be responded to quickly, fpw is reduced; when a long-term cascading fault needs to be captured, fpw is increased. Record each call and called microservice node, and integrate all the calling and called microservice nodes together. Finally, for each microservice node, obtain the number of calls initiated numCalling and the number of calls received numCalled of the microservice node. In addition, for the preset specific call , and the number of times it appears in the fault propagation time window is marked as count(a, b), where the preset specific call refers to the call that has significant characteristics within the fault propagation time window (fpw) and has an important impact on fault propagation, which is defined in advance.
[0069] Using all the above information, a fault impact graph skeleton is constructed to characterize the interaction of microservices and present all possible fault propagation paths during the fault propagation time window. In the fault impact graph, each microservice node is represented as a specific node, called Corresponding to the edges (a, b) and (b, a), (a, b) represents the fault propagation path from a to b, and (b, a) represents the fault propagation path from b to a. The skeleton of the fault impact graph is as follows: Figure 3 In the example shown, the microservice nodes in the fault impact diagram include: shopping cart, front-end, advertising, payment, checkout, and transportation. The lines between the microservice nodes represent the fault propagation path. The arrows in the fault propagation path show the fault propagation direction. The path segments used by the fault propagating from downstream to upstream are different from those used by the fault propagating from upstream to downstream. Figure 3 The fault propagation time window (fpw) in is set to 13.2 seconds.
[0070] Step S202: construct a fault impact graph skeleton based on each fault propagation path.
[0071] Among them, the fault impact diagram skeleton constructed based on each fault propagation path can be seen in Figure 3 shown.
[0072] Step S203: determining the fault propagation probability of each fault propagation path in the fault impact graph skeleton.
[0073] In the fault impact graph, the edge (a, b) represents the impact of a on b and reflects the fault propagation path from a to b. Among all the microservice nodes that may have an impact on b, different attributes of the microservice nodes will have different strengths of impact, thereby generating different fault propagation probabilities, which can be reflected in the different weights of the edges in the fault impact graph. That is, the present application uses the fault propagation probability from a to b as the weight of the edge corresponding to the fault propagation path from a to b in the fault impact graph.
[0074] Step S204: Add the fault propagation probability to the corresponding fault propagation path on the fault influence diagram skeleton to obtain a fault influence diagram.
[0075] It can be seen that the fault impact diagram in the present application not only shows the fault propagation path between microservice nodes, but also shows the fault propagation probability, thereby facilitating the subsequent root cause location of the fault.
[0076] In one embodiment, the process of determining the fault propagation probability of each fault propagation path in the fault impact graph skeleton may specifically include:
[0077] (1) Extract the fault propagation characteristics of each fault propagation path;
[0078] (2) The principal component analysis technique is used to process the fault propagation characteristics to obtain the fault propagation probability corresponding to the fault propagation path.
[0079] Among them, for each edge in the fault influence graph that serves as a fault propagation path, such as (a, b), the extracted fault propagation feature represents the degree of influence of a on b.
[0080] The fault propagation features in this application include but are not limited to node abnormality, node activity, intimacy and similarity between two microservice nodes in the fault propagation path.
[0081] Specifically, the node abnormality, node activity, intimacy and similarity between two microservice nodes in the fault propagation path are described in detail as follows:
[0082] (1) Node abnormality
[0083] Microservice node b is closely connected to microservice node a and is easily affected by the performance anomaly of microservice node a. Microservice node b usually has strong adaptability and can adapt to external interference, but the premise is that this influence is within the adaptive range. Therefore, the greater the degree of anomaly of microservice node a and the longer the impact lasts, the greater the possibility that microservice node b will be affected.
[0084] ADegree(a) is defined to measure the abnormality degree of microservice node a. The calculation method is shown in formula (2):
[0085] (2);
[0086] In the formula, Indicates the abnormality of microservice node a, s represents the constant factor, s is adjustable, and the value range is from 0 to 1. Represents the fault propagation time window Internal anomaly score The maximum value of represents the data point d detected at time t t The anomaly score, Indicates that in any fault propagation time window At each time point t within Indicates that in the fault propagation time window There is at least one time point t in Represents the fault propagation time window.
[0087] Among them, s is increased when the severity of abnormal node behavior is high or in high-risk scenarios, and vice versa.
[0088] (2) Node activity
[0089] Active microservice nodes interact frequently with other microservice nodes, act as bridges in information transmission, and act as agents in fault propagation, so they are more likely to affect other microservice nodes. The probability of microservice node a appearing in the shortest path between any two microservice nodes is used as a measure of the activity of microservice node a, as shown in formula (3):
[0090] (3);
[0091] In the formula, Indicates the activity of microservice node a, represents the total number of all pairs of shortest paths, represents the total number of paths passing through microservice node a, P represents microservice node P, q represents microservice node q, and service node a, microservice node P and microservice node q represent different microservice nodes.
[0092] (3) Intimacy
[0093] Intuitively, the more frequent the communication between microservice node a and microservice node b, the more likely microservice node a is to affect microservice node b. The intimacy of edge (a, b) is defined as shown in formula (4) to measure the intimacy between microservice node a and microservice node b. Formula (4) is as follows:
[0094] (4);
[0095] In the formula, Indicates call The proportion of communication in microservice node b, Indicates call The number of times (i.e. the number of times microservice node a and microservice node b communicate with each other), Indicates the number of communications of microservice node b (that is, the total number of times microservice node b is called).
[0096] (4) Similarity
[0097] If the performance indicators of two microservice nodes have similar change patterns, they may affect each other. Following this concept, the similarity of edge (a, b) is defined as shown in formula (5), which is as follows:
[0098] (5);
[0099] In the formula, Indicates the similarity between microservice node a and microservice node b, Represents the key business indicators of microservice node a. Represents the key business indicators of microservice node b. Represents the standard deviation of the key business indicators of microservice node a, Represents the standard deviation of the key business indicators of microservice node b, Represents the covariance of the key business indicators of microservice node a and microservice node b.
[0100] Among them, the similarity of (a, b) is derived from the Pearson correlation coefficient (a statistic used to measure the linear correlation between two variables), which is used to measure the correlation between the key business indicators of microservice node a and microservice node b. The similarity (a, b) is between -1 and 1, and the larger the absolute value, the higher the similarity.
[0101] This application uses principal component analysis technology to process fault propagation characteristics, and the process of obtaining the fault propagation probability corresponding to the fault propagation path is as follows:
[0102] For edge (a, b), the four fault propagation features shown above are combined to determine the impact of microservice node a on microservice node b and the probability of fault propagation from microservice node a to microservice node b. In order to combine all fault propagation features to calculate the final fault propagation probability, the present application uses principal component analysis technology to "compress" all fault propagation features, such as the four fault propagation features shown above, and the edge attributes contained therein into a single weight value w(a, b) almost losslessly, where the weight value w(a, b) represents the fault propagation probability from microservice node a to microservice node b, and the larger the value of w(a, b), the greater the fault propagation probability from microservice node a to microservice node b.
[0103] Principal Component Analysis (PCA) technology can transform high-dimensional data into low-dimensional representation, thereby simplifying the data model, improving computational efficiency, and facilitating data visualization and understanding.
[0104] In one embodiment, step S104 may specifically include:
[0105] Reverse the edges representing the fault propagation path in the fault influence graph to obtain a reverse fault influence graph;
[0106] The reverse fault impact graph is traversed along the reverse propagation path starting from the abnormal node to evaluate the suspicion score of each microservice node.
[0107] In practical applications, we can traverse the fault impact graph based on the exhaustive traversal method, starting from the abnormal node and traversing along the reverse propagation path, and simulate the inverse process of fault propagation through the suspicion score iteration method to evaluate the suspicion score of each microservice node.
[0108] Specifically, the fault impact graph represents the interaction between microservice nodes and models the fault propagation path of the entire microservice system. By reversing the edges in the fault impact graph, we can get an inverted fault impact graph (for ease of description, RIG is used to represent the inverted fault impact graph), where the edge (a, b) with weight w(a, b) indicates that microservice node a is affected by microservice node b, that is, the probability that the anomaly at microservice node a propagates from microservice node b is w(a, b).
[0109] Intuitively, for an abnormal node, if we recursively search for nodes that may cause the abnormality of the node along the edge of RIG, we can find the root cause of the fault. Based on this, this application adopts a reverse tracking algorithm starting from the abnormal node and exhaustively traversing RIG to track the root cause node that is most likely to cause the abnormal value.
[0110] This application assigns each microservice node a suspicion score (SS), which is accumulated by traversing the RIG as the probability that the microservice node corresponding to the suspicion score becomes the root cause node of the fault. The larger the suspicion score SS, the higher the probability that the microservice node corresponding to the suspicion score SS becomes the root cause node of the fault.
[0111] For microservice node b:
[0112] (1) The more links microservice node b receives in RIG, the more likely it is that microservice node b is the cause of the detected outlier, that is, the greater the possibility that b is the root cause node.
[0113] (2) In RIG, a link from a microservice node with a high suspicion score (SS) to microservice node b will increase the suspicion score of microservice node b because the possibility of microservice node b propagating the fault to other microservice nodes increases.
[0114] (3) The weight of the edge linked to microservice node b is also important because the weight represents the possibility of microservice node b being blamed.
[0115] Therefore, the suspicion score of microservice node b can be obtained by formula (6) and formula (7).
[0116] (6);
[0117] In the formula, represents the suspicion score of microservice node b, Indicates the suspicion score and weight value of microservice node a represents the original fault propagation probability from microservice node a to microservice node b, express The set of all edges in .
[0118] (7);
[0119] In the formula, represents the probability of fault propagation from microservice node a to microservice node b, Represents the fault transmission path from microservice node a to microservice node p, weight value Represents the slave microservice node The probability of fault propagation to microservice node p.
[0120] When performing a comprehensive evaluation of the suspicion score, a fault occurs on a microservice node and propagates in the microservice system through the fault propagation path, eventually leading to the detection of an abnormal node. In order to track the root cause node of the fault, the inverse process of fault propagation is simulated by exhaustively traversing the RIG and iterating formula (6) to comprehensively evaluate the node suspicion score. In practical applications, matrices can be used to speed up calculations. For a microservice system consisting of n microservice nodes, the corresponding RIG contains n microservice nodes. Let M be a transition matrix of n×n elements, each element M ba represents the probability that microservice node b causes a failure at microservice node a, which can be calculated by the formula shown in formula (8).
[0121] (8);
[0122] This application uses the vector SSV=[ss 0 ,ss 1 ,···, ] represents the set of suspicion scores of n microservice nodes, where ss 0 Indicates the suspicion score of the first microservice node, ss 1 Indicates the suspicion score of the second microservice node, Represents the suspicion score of the nth microservice node.
[0123] In addition, since the reverse tracking starts from the detected abnormal node, a seed vector Seed=[s 0 ,s 1 ,···, ],s 0 represents the initial suspicion score of the first microservice node, s 1 represents the initial suspicion score of the second microservice node, represents the initial suspicion score of the nth microservice node. If an outlier is detected in a microservice node, the microservice node corresponding to the anomaly is abnormal, and the corresponding element s i =1(i∈[0,n)).
[0124] The qth iteration of calculating the suspicion scores of n nodes is defined as:
[0125] (9);
[0126] Where M represents the n×n transition matrix, and each element M in the transition matrix ba Represents the probability that microservice node b causes a failure at microservice node a.
[0127] After the iterative calculation formula (9) converges, the suspicion scores of all microservice nodes can be obtained. All suspicion scores can be sorted in descending order, and the target microservice node with the highest suspicion score can be determined as the fault root cause node.
[0128] To understand the root cause location process of a fault in a microservice system, see Figure 4 , a technical flow chart of locating the root cause of a fault in a microservice system disclosed in an embodiment of the present application is as follows:
[0129] (1) Performance anomaly detection
[0130] The key business indicators of each microservice node in the microservice system are detected for performance anomalies. For example, the microservice system includes microservice nodes A, B, C, D, E, and F, and an abnormal value is detected in the key business indicator of microservice node B.
[0131] (2) Construct a fault impact diagram.
[0132] Specifically, it includes: extracting the fault propagation path from the calling relationship between microservice nodes within the fault propagation time window. For example, the microservice system includes microservice nodes A, B, C, D, E, and F. The fault transmission path is extracted from the calling relationship between microservice nodes A, B, C, D, E, and F, for example, including: (A, B), (B, D).
[0133] Extract the fault propagation features of (A, B), including feature 1 value, feature 2 value, feature 3 value and feature 4 value, and use principal component analysis technology to process these fault propagation features to obtain the corresponding fault propagation probability. This process is also called edge weight calculation.
[0134] The fault propagation features of (B, D) are extracted, including feature 1 value, feature 2 value, feature 3 value and feature 4 value, and these fault propagation features are processed using principal component analysis technology to obtain the corresponding fault propagation probability. This process is also called edge weight calculation.
[0135] (3) The microservice node corresponding to the outlier value is taken as the outlier node. Starting from the outlier node, the fault impact graph is traversed along the reverse propagation path, the suspicion score of each microservice node is evaluated, and the ranking is performed to obtain the ranking result. Assume that the ranking result is:
[0136] 1st place: B, 2nd place: A, 3rd place: D, 4th place: F, 5th place: C, 6th place: E.
[0137] Therefore, the microservice node B with the highest suspicion score is determined as the fault root cause node.
[0138] Corresponding to the above method embodiments, the present application also discloses a device for locating the root cause of a fault in a microservice system.
[0139] See also Figure 5 , a schematic diagram of the structure of a device for locating the root cause of a fault in a microservice system disclosed in an embodiment of the present application, the device may include:
[0140] The anomaly detection unit 301 is used to perform performance anomaly detection on key business indicators of each microservice node in the microservice system.
[0141] In a highly dynamic microservice environment, even if no anomalies occur, service performance will have peaks and valleys, so it is difficult to distinguish between true anomalies and normal fluctuations. Based on this, this application performs performance anomaly detection on the key business indicators of each microservice node in the microservice system to promptly detect abnormal values in key business indicators.
[0142] Preferably, the present application uses the isolation forest algorithm to detect performance anomalies for the key business indicators of each microservice node in the microservice system, thereby achieving high accuracy of anomaly detection. The isolation forest algorithm is an anomaly detection algorithm based on a decision tree, which detects outliers in a data set by constructing a set of randomly partitioned binary trees. The basic idea is to regard abnormal samples as isolated points in the feature space, while normal samples are densely distributed in the feature space.
[0143] It should be noted that detecting the key business indicators of microservice nodes is an effective means to evaluate the health status of microservice nodes and discover problems. The key business indicators of each microservice node may include one or more, see the example shown in Table 1.
[0144] The fault impact graph construction unit 302 is used to construct a fault impact graph simulating the fault propagation path between microservice nodes based on the abnormal value when an abnormal value is detected from all the key business indicators.
[0145] After detecting anomalies from key business indicators, this application will build a fault impact graph that simulates the fault propagation path between microservice nodes. The fault impact graph shows all possible fault propagation paths.
[0146] The score evaluation unit 303 is used to take the microservice node corresponding to the abnormal value as an abnormal node, traverse the fault impact graph along the reverse propagation path starting from the abnormal node, and evaluate the suspicion score of each microservice node.
[0147] In the fault impact graph, each node represents a microservice node, and the edge represents the fault propagation path between microservice nodes. Through the fault impact graph, we can directly understand the impact of a microservice node failure on other microservice nodes in the microservice system. This application takes the microservice node corresponding to the outlier as the abnormal node. Starting from the abnormal node, the process of tracing the fault source or the impact range along the reverse direction of the fault propagation path can be traced to determine the root cause node of the fault.
[0148] The suspicion score of a microservice node is a probability indicator used to measure the probability that a microservice node may have problems or failures in the microservice system. Starting from the abnormal node, while traversing the fault impact graph along the back propagation path, the suspicion score of each microservice node is evaluated to facilitate the subsequent determination of the final root cause node of the fault.
[0149] The root cause node determination unit 304 is configured to determine the target microservice node with the highest suspicion score as the fault root cause node.
[0150] The highest suspicion score indicates that the corresponding microservice node has a higher probability index of a problem or failure. Therefore, the present application determines the target microservice node with the highest suspicion score as the root cause node of the failure.
[0151] In summary, the present application discloses a device for locating the root cause of a fault in a microservice system, which performs performance anomaly detection on the key business indicators of each microservice node in the microservice system. When an outlier is detected from all the key business indicators, a fault impact graph simulating the fault propagation path between microservice nodes is constructed based on the outlier, and the microservice node corresponding to the outlier is taken as an abnormal node. The fault impact graph is traversed along the reverse propagation path from the abnormal node to evaluate the suspicion score of each microservice node, and the target microservice node with the highest suspicion score is determined as the root cause node of the fault. The present application simulates the fault propagation path between microservice nodes by constructing a fault impact graph, so as to intuitively present all potential fault propagation paths; by adopting reverse tracing technology, the entire fault impact graph is traversed along the reverse propagation path from the abnormal node to comprehensively evaluate the suspicion scores of each microservice node, so as to quickly screen out the microservice node with the highest suspicion score from all suspicion scores as the final root cause node of the fault. Therefore, the present application realizes the rapid and accurate positioning of the root cause node of the fault when a fault occurs.
[0152] In one embodiment, the anomaly detection unit 301 may be specifically used for:
[0153] The isolation forest algorithm is used to perform performance anomaly detection on the key business indicators of each microservice node in the microservice system.
[0154] In one embodiment, the fault impact graph construction unit 302 may be specifically used to:
[0155] When an abnormal value is detected from all the key business indicators, a fault propagation path is extracted from the call relationship between the microservice nodes within the fault propagation time window;
[0156] Constructing a fault impact graph skeleton based on each of the fault propagation paths;
[0157] Determining a fault propagation probability of each of the fault propagation paths in the fault impact graph skeleton;
[0158] The fault propagation probability is added to the corresponding fault propagation path on the fault impact diagram skeleton to obtain the fault impact diagram.
[0159] In one embodiment, the fault impact graph construction unit 302 may also be used to:
[0160] extracting a fault propagation feature of each of the fault propagation paths;
[0161] The fault propagation characteristics are processed using principal component analysis technology to obtain the fault propagation probability corresponding to the fault propagation path.
[0162] Among them, the fault propagation characteristics include: node abnormality, node activity, intimacy and similarity between two microservice nodes in the fault propagation path.
[0163] In one embodiment, the score evaluation unit 303 may be specifically configured to:
[0164] Reversing the edge representing the fault propagation path in the fault impact graph to obtain a reversed fault impact graph;
[0165] The reverse fault impact graph is traversed along a reverse propagation path starting from the abnormal node, and the suspicion score of each of the microservice nodes is evaluated.
[0166] In one embodiment, the score evaluation unit 303 may also be used to:
[0167] Based on the exhaustive traversal method, the inverse fault impact graph is traversed along the reverse propagation path starting from the abnormal node, and the inverse process of fault propagation is simulated by the suspicion score iteration method to evaluate and obtain the suspicion score of each microservice node.
[0168] It should be noted that, for the specific working principles of each component in the device embodiment, please refer to the corresponding part of the method embodiment, which will not be repeated here.
[0169] Corresponding to the above embodiments, the present application also discloses a computer storage medium, which stores at least one instruction. When the at least one instruction is executed by a processor, the steps shown in the embodiment of the method for locating the root cause of a fault in a microservice system are implemented.
[0170] Computer storage media may be tangible media that may contain or store programs for use by or in conjunction with an instruction execution system, device, or apparatus. Computer storage media may be machine-readable signal media or machine-readable storage media. Computer storage media may include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0171] Corresponding to the above embodiment, Figure 6 As shown, the present invention also provides an electronic device, which may include: a processor 1 and a memory 2;
[0172] The processor 1 and the memory 2 communicate with each other via a communication bus 3;
[0173] Processor 1, configured to execute at least one instruction;
[0174] Memory 2, used to store at least one instruction;
[0175] The processor 1 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention.
[0176] The memory 2 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0177] Among them, the processor executes at least one instruction to implement the steps shown in the embodiment of the method for locating the root cause of a fault in a microservice system.
[0178] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0179] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0180] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for locating the root cause of a fault in a microservice system, characterized in that: include: Perform performance anomaly detection on key business indicators of each microservice node in the microservice system; When an abnormal value is detected from all the key business indicators, a fault impact graph simulating a fault propagation path between microservice nodes is constructed based on the abnormal value; The microservice node corresponding to the abnormal value is taken as an abnormal node, and the fault impact graph is traversed along the reverse propagation path starting from the abnormal node to evaluate the suspicion score of each microservice node; The target microservice node with the highest suspicion score is identified as the fault root cause node.
2. The method for locating the root cause of a fault in a microservice system according to claim 1, characterized in that: The performance anomaly detection of key business indicators of each microservice node in the microservice system includes: The isolation forest algorithm is used to perform performance anomaly detection on the key business indicators of each microservice node in the microservice system.
3. The method for locating the root cause of a fault in a microservice system according to claim 1 or 2, characterized in that: When an abnormal value is detected from all the key business indicators, a fault impact graph simulating the fault propagation path between microservice nodes is constructed based on the abnormal value, including: When an abnormal value is detected from all the key business indicators, a fault propagation path is extracted from the call relationship between the microservice nodes within the fault propagation time window; Constructing a fault impact graph skeleton based on each of the fault propagation paths; Determining a fault propagation probability of each of the fault propagation paths in the fault impact graph skeleton; The fault propagation probability is added to the corresponding fault propagation path on the fault impact diagram skeleton to obtain the fault impact diagram.
4. The method for locating the root cause of a fault in a microservice system according to claim 3, characterized in that: The determining of the fault propagation probability of each fault propagation path in the fault impact graph skeleton comprises: extracting a fault propagation feature of each of the fault propagation paths; The fault propagation characteristics are processed using principal component analysis technology to obtain the fault propagation probability corresponding to the fault propagation path.
5. The method for locating the root cause of a fault in a microservice system according to claim 4, characterized in that: The fault propagation characteristics include: node abnormality, node activity, intimacy and similarity between two microservice nodes in the fault propagation path.
6. The method for locating the root cause of a fault in a microservice system according to claim 1, characterized in that: The step of taking the microservice node corresponding to the abnormal value as an abnormal node, traversing the fault impact graph along the reverse propagation path from the abnormal node, and evaluating the suspicion score of each microservice node includes: Reversing the edge representing the fault propagation path in the fault impact graph to obtain a reversed fault impact graph; The reverse fault impact graph is traversed along a reverse propagation path starting from the abnormal node, and the suspicion score of each of the microservice nodes is evaluated.
7. The method for locating the root cause of a fault in a microservice system according to claim 6, characterized in that: Starting from the abnormal node, traversing the reverse fault impact graph along the reverse propagation path, and evaluating the suspicion score of each microservice node, including: Based on the exhaustive traversal method, the inverse fault impact graph is traversed along the reverse propagation path starting from the abnormal node, and the inverse process of fault propagation is simulated by the suspicion score iteration method to evaluate and obtain the suspicion score of each microservice node.
8. A device for locating the root cause of a fault in a microservice system, characterized in that: include: The anomaly detection unit is used to detect performance anomalies of key business indicators of each microservice node in the microservice system; A fault impact graph construction unit is used to construct a fault impact graph simulating a fault propagation path between microservice nodes based on the abnormal value when an abnormal value is detected from all the key business indicators; A score evaluation unit, configured to take the microservice node corresponding to the abnormal value as an abnormal node, traverse the fault impact graph along a reverse propagation path starting from the abnormal node, and evaluate the suspicion score of each microservice node; The root cause node determination unit is used to determine the target microservice node with the highest suspicion score as the fault root cause node.
9. A computer storage medium, characterized in that The computer storage medium stores at least one instruction, and when the at least one instruction is executed by the processor, the method for locating the root cause of a fault in a microservice system according to any one of claims 1 to 7 is implemented.
10. An electronic device, characterized in that: The electronic device comprises: a memory and a processor; The memory is used to store at least one instruction; The processor is used to execute the at least one instruction to implement the method for locating the root cause of a fault in a microservice system as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Fault root cause positioning method and system for micro-service architecture information system
CN112698975A
Micro-service fault root cause positioning method based on fault propagation graph
CN114385397A
Microservice application system root cause positioning method and device, medium and equipment
CN114528175A
Cited By
Dynamic environment monitoring data visualization platform and abnormality diagnosis method
CN121030596A
Power environment monitoring data visualization platform and abnormal diagnosis method
CN121030596B
Method for analyzing abnormal root causes of distributed batch test results
CN121418266A
Distributed batch test result anomaly root cause analysis method
CN121418266B
Fault diagnosis method and device, storage medium and electronic equipment
CN121585525A