Abnormal root cause positioning method and device for micro service, medium and computer program product
By combining the abnormal indicators, graphs and update event text of microservices, using graph-text models and random walk algorithms, the problem of abnormal root cause positioning in the microservice architecture is solved, and fast and intelligent fault location and recovery are achieved.
Patent Information
- Application Number
- CN202510366191.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-01
AI Technical Summary
In the microservice architecture, there are great technical difficulties and challenges in how to quickly and intelligently locate the root cause of abnormalities and achieve failure recovery.
By combining the exception metrics of microservices, the graph of microservices and the update event text of microservices, the pre-trained graph-text model and the random walk algorithm, the correlation between the exception subgraph and the update event text is determined, and the weight is adjusted to prioritize the determination of the root cause node of the exception.
It realizes rapid positioning of the root cause of the fault, significantly shortens the troubleshooting time, and improves the system operation and maintenance efficiency.
Smart Images

Figure CN120238418A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to a method, device, medium, and computer program product for locating the root cause of exceptions in microservices. Background Art
[0002] In recent years, with the wide application of the microservices architecture, more and more enterprises have adopted a microservices-based system architecture to improve the flexibility and scalability of the system. However, the high complexity of the microservices architecture also brings new technical challenges. Especially when exceptions occur in the system, troubleshooting and locating errors usually require a large amount of R & D and operation and maintenance resources. This inefficient troubleshooting process may lead to a decline in the operational efficiency of enterprises and even cause significant economic losses. Therefore, in a microservices system, for exception scenarios, how to quickly and intelligently locate the root cause of errors and achieve fault recovery has become an important requirement in practical applications.
[0003] Existing technical solutions usually locate problems based on monitoring metrics, call chains, and logs. However, due to the distributed nature of the microservices architecture and the introduction of cloud-native technologies, how to effectively integrate these tools to quickly and accurately locate problems in a complex system still poses great technical difficulties and challenges. Summary of the Invention
[0004] The purpose of the present invention is to provide a method, device, medium, and computer program product for locating the root cause of exceptions in microservices, and to solve the technical problem of how to quickly locate the root cause of exceptions by combining exception metrics of microservices, graphs of microservices, and update event texts of microservices.
[0005] The first embodiment of the present invention discloses a method for locating the root cause of exceptions in microservices, which is used for an electronic device. The method includes:
[0006] Obtaining exception metrics of multiple microservices;
[0007] Obtaining graphs of the multiple microservices, where nodes of the graph represent microservices and edges represent call relationships existing between microservices;
[0008] Obtaining multiple update event texts of the multiple microservices;
[0009] Based on the exception metrics, assigning a first weight to the edges of the graph, where the first weight represents the degree of association between the nodes connected by the edge;
[0010] Based on a pre-trained graph-text model, determining the relevance between an exception subgraph and the multiple update event texts, where the exception subgraph of the graph is the weighted edges and the nodes connected by them;
[0011] In the case where the relevance reaches a predetermined threshold, for the updated nodes represented by the multiple update event texts in the abnormal subgraph, increase the second weight of the updated nodes, where the second weight represents the possibility that the node is preferentially determined as the abnormal root cause node in the random walk algorithm or the random teleportation algorithm;
[0012] Perform a random walk algorithm, or a random walk algorithm and a random teleportation algorithm, on the abnormal subgraph to determine the abnormal root cause node.
[0013] According to the first embodiment of the present invention, the obtaining of the abnormal metrics of multiple microservices includes:
[0014] Use the unsupervised anomaly detection algorithm of ADTK to obtain the microservice performance anomaly metrics and the microservice availability anomaly metrics, where the microservice performance anomaly metric is the response time of calls between microservices, and the microservice availability anomaly metric is the error count of microservices.
[0015] According to the first embodiment of the present invention, the assigning of the first weight to the edges of the graph based on the abnormal metrics includes:
[0016] According to the levels of the microservice performance anomaly metrics and the microservice availability anomaly metrics, use the Pearson correlation coefficient to assign the first weight to the edges between the nodes representing the abnormal microservices.
[0017] According to the first embodiment of the present invention, the obtaining of the graph of the multiple microservices includes:
[0018] Obtain the trace data within a predetermined time period, including the call relationship and call duration between microservices within the predetermined time period;
[0019] Generate the edges of the graph according to the trace data;
[0020] Obtain the description texts of the multiple microservices.
[0021] According to the first embodiment of the present invention, the determining of the relevance between the abnormal subgraph and the multiple update event texts based on the pre-trained graph-text model includes:
[0022] Extract the first vector of the description text of the microservice corresponding to the node in the abnormal subgraph as the attribute of the node;
[0023] Extract the second vector of the update event text corresponding to the node in the multiple update event texts;
[0024] Use the graph attention network to extract the features of the first vector;
[0025] Determine the similarity between the feature and the second vector.
[0026] According to the first embodiment of the present invention, the first vector and the second vector are extracted by using the BERT model.
[0027] According to the first embodiment of the present invention, the updated event text includes the description text of the updated api interface.
[0028] According to the first embodiment of the present invention, it further includes:
[0029] When the relevance reaches a predetermined threshold, a virtual node is newly created. The virtual node points to the updated node represented by the multiple updated event texts in the abnormal subgraph, and the second weight of the virtual node is 1.
[0030] According to the first embodiment of the present invention, performing a random walk algorithm, or a random walk algorithm and a random teleportation algorithm on the abnormal subgraph to determine the abnormal root cause node includes:
[0031] Performing the Personalized PageRank algorithm on the abnormal subgraph to determine the abnormal root cause node.
[0032] The second embodiment of the present invention discloses an electronic device, which includes a memory storing computer-executable instructions and a processor. When the instructions are executed by the processor, the electronic device implements the method for locating the abnormal root cause of the microservice according to the first embodiment of the present invention.
[0033] The third embodiment of the present invention discloses a computer storage medium, on which instructions are stored. When the instructions run on a computer, the computer executes the method for locating the abnormal root cause of the microservice according to the first embodiment of the present invention.
[0034] The fourth embodiment of the present invention discloses a computer program product, including computer-executable instructions, and the instructions are executed by a processor to implement the method for locating the abnormal root cause of the microservice according to the first embodiment of the present invention.
[0035] Compared with the prior art, the main differences and effects of the embodiments of the present invention are as follows:
[0036] Through the technical solution of the present invention, by combining monitoring metrics, microservice graph data including call relationships, and the relevance of update event texts, rapid location of the root cause of a fault is achieved. Among them, an abnormal subgraph is obtained by combining monitoring metrics and microservice graph data; then, by combining the update event text, the weight of the abnormal subgraph is adjusted, taking into account the association between the update event and the abnormality; finally, random walk, or random walk and random teleportation, is performed on the abnormal subgraph with adjusted weights to extract the dynamic and global information of the abnormal subgraph and find the root cause node of the abnormality. The root cause microservice corresponding to the abnormality can be quickly determined, thereby significantly shortening the troubleshooting time and improving the system operation and maintenance efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 The flowchart showing the method for locating the root cause of an abnormality in a microservice according to an embodiment of the present application.
[0038] Figure 2 The microservice abnormality root cause location system shown according to an embodiment of the present application.
[0039] Figure 3 The hardware structure block diagram of an electronic device for locating the root cause of an abnormality in a microservice shown according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0040] To make the objectives, technical solutions, and advantages of the present application clearer, the embodiments of the present application will be described in further detail below in conjunction with the accompanying drawings.
[0041] In view of the technical problem that the prior art cannot quickly locate the root cause of an abnormality by combining the abnormal metrics of microservices, the graph of microservices, and the update event texts of microservices, the present application provides a method for locating the root cause of an abnormality in a microservice for an electronic device. Referring to the flowchart as shown in Figure 1 The method includes:
[0042] S101, obtaining abnormal metrics of multiple microservices; obtaining graphs of multiple microservices, where the nodes of the graph represent microservices and the edges represent the call relationships existing between microservices; obtaining multiple update event texts of multiple microservices.
[0043] According to the first embodiment of the present invention, obtaining graphs of multiple microservices includes: obtaining trace data within a predetermined time period, including the call relationships and call durations between microservices within the predetermined time period; generating the edges of the graph according to the trace data; obtaining description texts of multiple microservices.
[0044] Figure 2 The microservice abnormality root cause location system shown according to an embodiment of the present application includes a microservice system 201, an anomaly detection module 202, a call link graph construction module 203, and an API dynamic document storage module 204.
[0045] The microservice system 201 is a system that contains multiple microservices, and the multiple microservices call each other through API interfaces.
[0046] The anomaly detection module 202 is the starting point for locating the root cause of microservice failures. This module uses the unsupervised anomaly detection algorithm of ADTK to detect spikes and level changes in time series. Its principle is to scan a time series through two parallel time windows, generate a new time series through the median difference between the two time windows. Since the overall time range is short, the transformed series should conform to the definition of a stationary time series. A "stationary" time series is statistically defined as having statistical characteristics independent of the observed time points. For this transformed time series, the outlier interval is selected by setting parameters using the interquartile range method. The interquartile range method means first calculating IQR = Q3 - Q1, where Q3 is the third quartile and Q1 is the first quartile. This interquartile range method defines an outlier interval [Q1 - c×IQR, Q3 + c×IQR], where c is a constant used to adjust the sensitivity of anomaly detection. If the points in the transformed series exceed this range, they are considered outliers. The advantages of using this interquartile range method include that it is less affected by extreme values as it only considers the median value of the data, and it does not require the assumption of a normal distribution. Additionally, since the final root cause reasoning module will use the method of calculating the time series based on the Pearson correlation coefficient for root cause reasoning, the recall rate of anomaly detection is more important than the accuracy rate in this scenario, and this interquartile range method can ensure a high recall rate. In this application, the anomaly metrics collected by the anomaly detection module 202 include microservice performance anomalies and microservice availability anomalies.
[0047] Microservice performance anomalies use the response time of calls between microservices as a reference metric. For example, in an actual business production environment, the response time of microservices is generally less than 100 ms. When an anomaly occurs, the response time often exceeds 1 s. The reasons for microservice response time anomalies may include application thread blocking, high load, etc. Microservice availability anomalies use the error count as their measurement metric, and the error count of normally running microservices is very small. For example, in an actual production environment, 15 times per second can be set as the threshold for microservice availability anomalies to occur.
[0048] The call link graph construction module 203 can construct a link graph of the call links (call relationships) between microservices in the microservice system 201, with the aim of subsequently constructing an exception propagation subgraph when a microservice exception occurs. When an error occurs, the error often propagates according to the call relationships between microservices. Therefore, the topological relationship of the services is the main basis for checking the root cause of the failure.
[0049] In a distributed link tracing system, client components compliant with the OpenTelemetry standard are installed in each microservice. These components enable the application to record the call relationships and call durations of services when inter-service calls occur. These trace information is assembled through spanid and traceid and stored in a time series database. The online call link graph construction module 203 obtains the call relationship graph between services through trace data over a period of time. The "graph" in this article refers to "graph data", which is a structured data form representing entities and their relationships using nodes (points) and edges (connections), and is often used to describe complex association networks. If microservice A makes a call to microservice B, there is an edge starting from node A and pointing to node B. The online call link graph construction module 203 will generate an online microservice call relationship graph based on the call link data within a certain period of time (such as within 30 minutes) after the exception occurs. It reflects the call relationships between each microservice in the past 30 minutes. It should be noted that there may be many API calls from microservice A to microservice B. In this module, they are aggregated without considering more detailed edges. The Pearson correlation coefficient is used to assign weights to the edges between the faulty microservice nodes.
[0050] The API documentation is a corpus used for offline training of the Graph-Text model. Since the API documentation itself includes information on both the link graph and the text, using the data in the API documentation to train the Graph-Text model enables the Graph-Text model to learn the correlation between business text and the microservice link graph. The main function of the API dynamic documentation storage module 204 is to store and dynamically manage the API documentation of all microservices online. The API documentation of microservices is an information document that describes the API interfaces provided by microservices, including the endpoints, request methods, parameters, response formats, and examples of each API. Well-maintained API documentation can improve development efficiency, reduce development time, lower communication costs, and provide important reference for automated testing and microservice fault location. In the entire microservice system 201, for example, in devops, the Swagger documentation is used to define the APIs in a standardized way, and while following the OpenAPI specification, clear text descriptions are given for all API interfaces. This module connects the large-scale microservice system 201 running online and the static domain knowledge of the entire microservice system 201. The most important aspect is to match the url path of the istio log with the url of the API to obtain the text domain meaning of the log. To keep the domain knowledge of this module updated, the implementation of this module is integrated into the pipeline of the entire system (such as devops). That is, when a specific interface of a microservice application is newly added or changed, the script of the build pipeline will perform code scanning and update the content of the API documentation library, and ensure that all API interfaces running online can find the descriptive text information of their text meanings.
[0051] S102, Based on the anomaly metrics, assign a first weight to the edges of the graph, where the first weight represents the degree of association between the nodes connected by the edge.
[0052] For example, the microservice root cause of anomaly location system further includes a graph structure processing module 205. Based on the anomaly metrics obtained from the anomaly detection module 202 and the link graph of the call relationships between microservices in the microservice system 201 constructed by the call link graph construction module 203, assign a first weight to the edges of the graph. Specifically, it may include: According to the levels of the microservice performance anomaly metrics and the microservice availability anomaly metrics, use the Pearson correlation coefficient to assign values to the first weights of the edges between the nodes representing the anomaly microservices.
[0053] S103, Based on the pre-trained graph-text model, determine the relevance between the anomaly subgraph and multiple update event texts, where the anomaly subgraph of the graph is the weighted edges and the nodes they connect;
[0054] According to the first embodiment of the present invention, based on a pre-trained graph-text model, determining the relevance between an abnormal subgraph and multiple update event texts includes: extracting a first vector of the description text of the microservice corresponding to the node in the abnormal subgraph as the attribute of the node; extracting a second vector of the update event text corresponding to the node in the multiple update event texts; using Graph Attention Networks (GAT) to extract the features of the first vector; and determining the similarity between the features and the second vector.
[0055] For example, the microservice abnormal root cause location system further includes a Graph-Text model 207 (graph-text model), and the Graph-Text model is a binary classification model trained offline. Its main task is to determine whether there is a correlation between an abnormal propagation graph and an external text. In the prediction stage, the input of the Graph-Text model is the online error propagation graph, the basic information of the node, and the text updated by the external microservice system 201. In the devops system, when the system development engineer enters the update text, it is clearly marked which api function changes are involved in this update and what changes have occurred to the api function. The Graph-Text model combines the data of the graph structure and the data of the text structure and finally gives the correlation between the two. During the training process, the call link data in the production environment and the description data of its api documentation are used for training. For the data of the graph structure, a graph attention neural network can be used for feature extraction, and for the data of the text structure, a pre-trained Bert model can be used to extract the features of the text content. The specific method is as follows.
[0056] Assume that the current microservice exception propagation graph is G(N, E), where nodei (0 ≤ i < N) represents the microservice node with an exception, and edgej (0 ≤ j < E) represents the call edge between the microservice nodes with exceptions. For each node Ni, let xi represent the natural language text obtained by merging the name of the current microservice node and the microservice node description. Use the pre-trained Bert model 1 to extract the hidden layer vector of the pooling layer output of xi. For each xi, i ∈ [0, N), hi can be obtained as the attribute of the node itself. Its vector shape is (1, F), where F = 768. Similarly, assume that the text of the most recent update event of the current exception node nodei is denoted as yi, and use the same text processing method to obtain βi as its Bert hidden layer vector, whose shape is also (1, F), where F = 768. Here, only the text of the microservice node is used as the node attribute (instead of the specific API text information on this link), because it is assumed that many links in the current system have failed. The reason may be that a global thread blockage at the application level has occurred in a certain node. In this extreme case, most link alarms may be obtained. Engineers need to view the problem from a global perspective. Therefore, the error propagation graph at the microservice node dimension is considered, rather than the error propagation graph at the specific API level.
[0057] Then, use two Graph Attention Layers to extract the features of the node attributes. Set F′ = 768 as the number of hidden layer features output by the first Graph Attention Layer, and F″ = 768 as the number of hidden layer features output by the second Graph Attention Layer.
[0058] The graph neural network then needs to calculate the self-attention scores. Here, the self-attention scores only need to be calculated between the nodes pointing to the current nodei and the nodei itself. Its meaning is the message passing mechanism of the graph neural network, that is, the vector representation of a node is jointly determined by its neighboring nodes pointing to it and itself, and the importance they determine is different. Let eij represent the self-attention coefficient, which represents the importance of nodej to nodei, and its value is calculated by formula (1), where a is a combination of a simple linear function and a ReLU non-linear activation function, which transforms two F′-dimensional vectors into a scalar as the attention weight coefficient. Here, a is the shared attention parameter learned in the graph neural network. The fully connected weight matrix W includes learnable parameters, and all nodes in the current graph share the weights of W.
[0059]
[0060] Furthermore, in order to make the attention scores of the incoming edge nodes and the current node comparable, use the softmax function to regularize the coefficients of these nodes, and calculate using formula (2), where represents the set of incoming edge nodes of the current node \(i\), which has \(j\) node elements.
[0061]
[0062] After considering the information of the nodes around node \(i\), the features of node \(i\) are used to extract the information of neighbor nodes and itself according to formula (3) with weights, and then through a non-linear ReLU function, the transformed \(h_i'\) is obtained, which is the final output of the graph attention layer, with a shape of \((1, F')\).
[0063]
[0064] In addition, to improve the interpretability of the model and avoid overfitting, the algorithm in this paper uses multi-head attention with \(K = 2\) to join the GAT layer, and takes the mean after adding the vectors of different layers. The corresponding formula is formula (4). Similar to Transformers, in order to stabilize the training, layer normalization is used to regularize the feature dimensions of each node to prevent the problem of gradient disappearance or gradient explosion that may occur during the training process. Using layer normalization for nodes eliminates the dependence on the node feature distribution and scale, so it can improve the generalization ability of the model for different node features.
[0065]
[0066] Finally, all \(h_i'\) are integrated as shown in formula (5), and further the similarity with \(\beta\) is calculated using the sigmoid function as shown in formula (6).
[0067]
[0068] \(y = Sigmoid(WConcat(g, \beta)+b)\) (6)
[0069] S104, in the case where the relevance reaches a predetermined threshold, for the updated nodes represented by multiple update event texts in the abnormal subgraph, increase the second weight of the updated nodes, and the second weight represents the possibility that the node is preferentially determined as the abnormal root cause node in the random walk algorithm or the random teleportation algorithm;
[0070] According to the first embodiment of the present invention, it further includes: in the case where the relevance reaches a predetermined threshold, create a virtual node, the virtual node points to the updated nodes represented by multiple update event texts in the abnormal subgraph, and the second weight of the virtual node is 1.
[0071] S105, perform a random walk algorithm, or a random walk algorithm and a random teleportation algorithm on the abnormal subgraph to determine the abnormal root cause node.
[0072] For example, the microservice exception root cause localization system further includes a root cause reasoning module 208. The root cause reasoning module 208 mainly uses the unsupervised learning algorithm of Personalized PageRank (PPR) to perform random walks on the error propagation graph and outputs the root cause. Its core idea is to determine the probability of moving to the next node based on the correlation of abnormal indicators between microservices. On this basis, based on the output of the Graph-Text model 207, after determining that the similarity between the relevant release text and the current error propagation graph reaches a certain threshold, a virtual node is created. The virtual node points to the currently updated microservice and is set with a weight of 1. Since these virtual nodes have no incoming edges, they cannot be accessed through random walks themselves and can only be accessed through random teleportation. Therefore, they will not ultimately become the root cause of the failure. By creating virtual nodes, the number of times the updated microservice is accessed during random walks is increased. However, assume that the update text of a microservice is highly relevant to the online link, but it is not on the microservice error propagation graph at all and is a normal or isolated abnormal node. In this case, its weight will still not be increased. Thus, it naturally achieves a double-verification effect. The root cause reasoning module 208 mainly traverses the microservice error propagation graph based on the similarity of errors between microservices and also considering the update text of external microservices to find the root cause of the failure.
[0073] Here, when the graph structure processing module 205 obtains the abnormal propagation subgraph of the anomaly detection module 202 and has obtained the index anomaly correlation coefficient of the connected microservices, a weighted abnormal propagation graph is obtained. After applying the PPR algorithm to this abnormal propagation graph, the root cause ranking under the condition of only considering index correlation can be obtained. In the PPR algorithm, the finally obtained PageRank vector (PPV) v is used, and v represents the "abnormal score" of each microservice node becoming the root cause of the failure. To calculate the PPV, first, the state transition matrix M of the PageRank algorithm needs to be defined. The calculation of the state transition matrix is shown in formula (7). Different from the equal importance of each node in the traditional PageRank algorithm, in this problem, the correlation coefficient of the indexes between nodes is used to assign values to M.
[0074]
[0075] That is, the higher the correlation of the abnormal indicators, the greater the probability of random walk access between them. Then, the state update equation of PPR is defined as formula (8), where u is the personalization vector, that is, the second weight, which gives nodes with higher anomaly detection confidence more access opportunities, and β is the damping coefficient of PPR, which represents the probability of moving along the edge or random teleportation.
[0076] υ = βMυ+(1 - β)u (8)
[0077] Through the technical solution of the present invention, by combining monitoring metrics, microservice graph data including call relationships, and the text relevance of update events, the rapid location of the root cause of a fault is achieved. Among them, an abnormal subgraph is obtained by combining monitoring metrics and microservice graph data; then, by combining the update event text, the weight of the abnormal subgraph is adjusted, and the correlation between the update event and the abnormality is taken into account; finally, random walk, or random walk and random teleportation, is performed on the abnormal subgraph with adjusted weights to extract the dynamic and global information of the abnormal subgraph and find the root cause node of the abnormality. The root cause microservice corresponding to the abnormality can be quickly determined, thereby significantly shortening the fault troubleshooting time and improving the system operation and maintenance efficiency.
[0078] The method embodiment of the present application corresponds to this embodiment, and this embodiment can be implemented in cooperation with the method embodiment of the present application. The relevant technical details mentioned in the method embodiment of the present application are still valid in this embodiment. To avoid repetition, they are not elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the method embodiment of the present application.
[0079] Figure 3 It is a hardware structure block diagram of an electronic device for implementing the root cause location of microservice abnormalities according to the embodiments of the present application.
[0080] As Figure 3 shown, the electronic device 300 may include one or more processors 302, a system motherboard 308 connected to at least one of the processors 302, a system memory 304 connected to the system motherboard 308, a non-volatile memory (NVM) 306 connected to the system motherboard 308, and a network interface 310 connected to the system motherboard 308.
[0081] The processor 302 may include one or more single-core or multi-core processors. The processor 302 may include any combination of general-purpose processors and dedicated processors (such as graphics processors, application processors, baseband processors, etc.). In the embodiments of the present invention, the processor 302 may be configured to execute one or more embodiments according to the method embodiments of the present application.
[0082] In some embodiments, the system motherboard 308 may include any suitable interface controller to provide any suitable interface to at least one of the processors 302 and / or any suitable device or component communicating with the system motherboard 308.
[0083] In some embodiments, the system motherboard 308 may include one or more memory controllers to provide an interface to the system memory 304. The system memory 304 may be used to load and store data and / or instructions. In some embodiments, the system memory 304 of the electronic device 300 may include any suitable volatile memory, such as a suitable dynamic random access memory (DRAM).
[0084] The NVM 306 may include one or more tangible, non-transitory computer-readable media for storing data and / or instructions. In some embodiments, the NVM 306 may include any suitable non-volatile memory such as flash memory and / or any suitable non-volatile storage device, such as at least one of a hard disk drive (HDD), a compact disc (CD) drive, and a digital versatile disc (DVD) drive.
[0085] The NVM 306 may include a portion of the storage resources installed on the device of the electronic device 300, or it may be accessible by the device but not necessarily part of the device. For example, the NVM 306 may be accessed via the network interface 310 over a network.
[0086] Specifically, the system memory 304 and the NVM 306 may respectively include: a temporary copy and a permanent copy of the instructions 320. The instructions 320 may include: instructions that, when executed by at least one of the processors 302, cause the electronic device 300 to implement the methods of the present application. In some embodiments, the instructions 320, hardware, firmware, and / or its software components may alternatively be disposed in the system motherboard 308, the network interface 310, and / or the processor 302.
[0087] The network interface 310 may include a transceiver for providing a radio interface for the electronic device 300 to communicate with any other suitable device (such as a front-end module, an antenna, etc.) over one or more networks. In some embodiments, the network interface 310 may be integrated with other components of the electronic device 300. For example, the network interface 310 may be integrated with at least one of the processor 302, the system memory 304, the NVM 306, and a firmware device with instructions (not shown), and when at least one of the processors 302 executes the instructions, the electronic device 300 implements one or more embodiments of the various method embodiments of the present application.
[0088] The network interface 310 may further include any suitable hardware and / or firmware to provide a multiple-input multiple-output radio interface. For example, the network interface 310 may be a network adapter, a wireless network adapter, a telephone modem, and / or a wireless modem.
[0089] In one embodiment, at least one of the processors 302 may be packaged with one or more controllers for the system motherboard 308 to form a System in Package (SiP). In one embodiment, at least one of the processors 302 may be integrated with one or more controllers for the system motherboard 308 on the same die to form a System on Chip (SoC).
[0090] The electronic device 300 may further include: an Input / Output (I / O) device 312, connected to the system motherboard 308. The I / O device 312 may include a user interface that enables a user to interact with the electronic device 300; the design of the peripheral component interface enables peripheral components to also interact with the electronic device 300. In some embodiments, the electronic device 300 further includes sensors for determining at least one of environmental conditions and location information related to the electronic device 300.
[0091] In some embodiments, the I / O device 312 may include, but is not limited to, a display (e.g., a liquid crystal display, a touch screen display, etc.), a speaker, a microphone, one or more cameras (e.g., a still image camera and / or a video camera), a flashlight (e.g., a light emitting diode flash), and a keyboard.
[0092] In some embodiments, the peripheral component interface may include, but is not limited to, a non-volatile memory port, an audio jack, and a power interface.
[0093] In some embodiments, the sensors may include, but are not limited to, a gyroscope sensor, an accelerometer, a proximity sensor, an ambient light sensor, and a positioning unit. The positioning unit may also be part of or interact with the network interface 310 to communicate with components of a positioning network (e.g., Global Positioning System (GPS) satellites).
[0094] It can be understood that the structure illustrated in the embodiments of the present invention does not constitute a specific limitation on the electronic device 300. In other embodiments of the present application, the electronic device 300 may include more or fewer components than those illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0095] Program code may be applied to input instructions to perform the various functions described in the present invention and generate output information. The output information may be applied to one or more output devices in a known manner. For the purposes of the present application, a system for processing instructions including the processor 302 includes any system having a processor such as a Digital Signal Processor (DSP), a microcontroller, an Application Specific Integrated Circuit (ASIC), or a microprocessor.
[0096] The program code can be implemented in a high-level programming language or an object-oriented programming language to communicate with the processing system. When necessary, the program code can also be implemented in an assembly language or a machine language. In fact, the mechanisms described in the present invention are not limited to the scope of any specific programming language. In any case, the language can be a compiled language or an interpreted language.
[0097] One or more aspects of at least one embodiment can be implemented by instructions stored on a computer-readable storage medium, which, when read and executed by a processor, enable an electronic device to implement the methods of the embodiments described in the present invention.
[0098] According to some embodiments of the present application, a computer storage medium is disclosed, on which instructions are stored, and when the instructions run on a computer, the computer is enabled to execute the method for locating the root cause of exceptions in microservices according to the embodiments of the present application.
[0099] The method embodiments of the present application correspond to this embodiment, and this embodiment can be implemented in cooperation with the method embodiments of the present application. The relevant technical details mentioned in the method embodiments of the present application are still valid in this embodiment. To avoid repetition, they are not elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the method embodiments of the present application.
[0100] According to some embodiments of the present application, a computer program product is disclosed, including computer-executable instructions that, when executed by a processor, implement the method for locating the root cause of exceptions in microservices according to the embodiments of the present application.
[0101] The method embodiments of the present application correspond to this embodiment, and this embodiment can be implemented in cooperation with the method embodiments of the present application. The relevant technical details mentioned in the method embodiments of the present application are still valid in this embodiment. To avoid repetition, they are not elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the method embodiments of the present application.
[0102] It can be understood that the specific embodiments described herein are merely for explaining the present application and not for limiting the present application. In addition, for the sake of convenience of description, only parts related to the present application rather than all structures or processes are shown in the drawings. It should be noted that in this specification, similar reference numerals and letters denote similar items in the drawings.
[0103] It should be understood that although terms such as "first", "second", etc. may be used herein to describe various features, these features should not be limited by these terms. These terms are used only for distinction and should not be construed as indicating or implying relative importance. For example, without departing from the scope of the exemplary embodiments, the first feature may be referred to as the second feature, and similarly, the second feature may be referred to as the first feature.
[0104] In the description of the present application, it should also be noted that unless otherwise clearly specified and defined, the terms "arranged", "connected", and "coupled" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present embodiment can be understood according to specific circumstances.
[0105] The illustrative embodiments of the present application include, but are not limited to, methods, devices, media, and computer program products for locating the root cause of anomalies in microservices.
[0106] The various aspects of the illustrative embodiments will be described using terms commonly employed by those skilled in the art to convey the substance of their work to others skilled in the art. However, it will be apparent to those skilled in the art that some alternative embodiments can be implemented using some of the described features. For purposes of explanation, specific numbers and configurations are set forth to provide a more thorough understanding of the illustrative embodiments. However, it will be apparent to those skilled in the art that alternative embodiments can be implemented without specific details. In some other cases, some well-known features are omitted or simplified to avoid obscuring the illustrative embodiments of the present application.
[0107] In addition, the various operations will be described as a plurality of operations that are separated from each other in a manner most conducive to understanding the illustrative embodiments; however, the described order should not be construed as implying that these operations must be dependent on the described order, and many of these operations can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the operations can also be rearranged. When the described operations are completed, the process can be terminated, but there may also be additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0108] References in the specification to "one embodiment", "an embodiment", "an illustrative embodiment", etc., mean that the described embodiment may include a particular feature, structure, or characteristic, but every embodiment may or may not necessarily include the particular feature, structure, or characteristic. Moreover, these phrases are not necessarily referring to the same embodiment. Further, when a particular feature is described in connection with a specific embodiment, the knowledge of those skilled in the art can affect the combination of such feature with other embodiments, whether or not such embodiments are explicitly described.
[0109] Unless the context otherwise requires, the terms "comprise", "have", and "include" are synonyms. The phrase "A and / or B" means "(A), (B), or (A and B)".
[0110] As used herein, the term "module" may refer to, as part of, or include: a memory (shared, dedicated, or group) for running one or more software or firmware programs, an application specific integrated circuit (ASIC), an electronic circuit, and / or a processor (shared, dedicated, or group), combinational logic circuitry, and / or other suitable components that provide the described functionality.
[0111] In the drawings, some structural or method features may be shown in a particular arrangement and / or order. However, it should be understood that such particular arrangement and / or ordering is not required. Rather, in some embodiments, these features may be illustrated in a different manner and / or order than shown in the illustrative drawings. Additionally, the structural or method features included in a particular drawing do not mean that all embodiments need to include such features. In some embodiments, these features may be omitted or may be combined with other features.
[0112] In some cases, the disclosed embodiments may be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments may also be implemented in the form of instructions or programs carried or stored on one or more transient or non-transient machine-readable (e.g., computer-readable) storage media, which may be read and executed by one or more processors and the like. When the instructions or programs are run by a machine, the machine may perform the various methods described above. For example, the instructions may be distributed via a network or other computer-readable media. Thus, machine-readable media may include, but are not limited to, any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), such as a floppy disk, an optical disk, a compact disc read-only memory (CD-ROMs), a magneto-optical disk, a read-only memory (ROM), a random access memory (RAM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic or optical card, or a flash memory or tangible machine-readable memory for transmitting network information by electrical, optical, acoustic, or other forms of signals (e.g., carrier waves, infrared signals, digital signals, etc.). Thus, machine-readable media include any form of machine-readable media suitable for storing or transmitting electronic instructions or machine (e.g., computer) readable information.
[0113] The embodiments of the present application have been described in detail above in conjunction with the accompanying drawings. However, the use of the technical solutions of the present application is not limited to the various applications mentioned in the embodiments of the present application. Various structures and variations can be easily implemented with reference to the technical solutions of the present application to achieve the various beneficial effects mentioned herein. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes made without departing from the spirit of the present application shall fall within the scope covered by the patent of the present application.
Claims
1. A method for locating the root cause of anomalies in a microservice, used in an electronic device, characterized in that: The method comprises: Get abnormal indicators of multiple microservices; Obtain a graph of the multiple microservices, wherein the nodes of the graph represent microservices, and the edges represent call relationships between the microservices; Obtain multiple update event texts of the multiple microservices; Based on the abnormality indicator, assign a first weight to the edge of the graph, wherein the first weight represents the degree of association between nodes connected by the edge; Based on a pre-trained graph-text model, determining the relevance of an abnormal subgraph to the multiple update event texts, wherein the abnormal subgraph of the graph is a weighted edge in the graph and the nodes connected thereto; When the correlation reaches a predetermined threshold, for the update node represented by the multiple update event texts in the abnormal subgraph, increase a second weight of the update node, wherein the second weight represents the possibility of the node being preferentially determined as the abnormal root cause node in the random walk algorithm or the random transmission algorithm; A random walk algorithm, or a random walk algorithm and a random transmission algorithm are executed on the abnormal subgraph to determine the abnormal root cause node.
2. The method according to claim 1, characterized in that: The abnormal indicators of multiple microservices are obtained, including: Use ADTK's unsupervised anomaly detection algorithm to obtain microservice performance anomaly indicators and microservice availability anomaly indicators, where the microservice performance anomaly indicator is the response time of calls between microservices, and the microservice availability anomaly indicator is the error count of the microservice.
3. The method according to claim 2, characterized in that The assigning a first weight to the edge of the graph based on the abnormal indicator includes: According to the level of the microservice performance abnormality index and the microservice availability abnormality index, the Pearson correlation coefficient is used to assign the first weight of the edge between the nodes representing the abnormal microservice.
4. The method according to claim 1, characterized in that: The obtaining of the graphs of the multiple microservices includes: Obtain trace data within a predetermined time period, including the call relationship and call duration between microservices within the predetermined time period; Generate edges of the graph according to the trace data; Get description texts of the multiple microservices.
5. The method according to claim 4, characterized in that The determining, based on the pre-trained graph-text model, the relevance between the abnormal subgraph and the plurality of update event texts comprises: Extracting a first vector of the description text of the microservice corresponding to the node in the abnormal subgraph as an attribute of the node; Extracting a second vector of update event text corresponding to the node from the plurality of update event texts; Using a graph attention network, extracting features of the first vector; A similarity between the feature and the second vector is determined.
6. The method according to claim 5, characterized in that The first vector and the second vector are extracted using a BERT model.
7. The method according to claim 1, characterized in that The update event text includes a description text of the updated API interface.
8. The method according to claim 1, characterized in that Also includes: When the correlation reaches a predetermined threshold, a new virtual node is created, the virtual node points to the update node represented by the multiple update event texts in the abnormal subgraph, and the second weight of the virtual node is 1.
9. The method according to claim 1, characterized in that: The performing of a random walk algorithm, or a random walk algorithm and a random transmission algorithm on the abnormal subgraph to determine the abnormal root cause node includes: A Personalized PageRank algorithm is executed on the abnormal subgraph to determine the abnormal root cause node.
10. An electronic device, characterized in that: The electronic device comprises a memory storing computer executable instructions and a processor. When the instructions are executed by the processor, the electronic device implements the method for locating the abnormal root cause of the microservice according to any one of claims 1 to 9.
11. A computer storage medium, characterized in that: Instructions are stored on the computer storage medium. When the instructions are executed on a computer, the computer is enabled to execute the method for locating the abnormal root cause of a microservice according to any one of claims 1 to 9.
12. A computer program product, characterized in that The method comprises computer executable instructions, wherein the instructions are executed by a processor to implement the method for locating the abnormal root cause of a microservice according to any one of claims 1 to 9.