A microservice fault root cause localization method, system, and device
Patent Information
- Application Number
- CN202610796379.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-04
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-06-04
AI Technical Summary
[0006]综上所述,现有技术在以下方面仍存在不足:信息表征的片面性(难以深度融合多模态数据)、粒度视角的单一性(忽略全局拓扑特征与局部异常的协同)以及故障上下文的孤立性(缺乏有效的跨故障样本知识迁移能力)
[0022] First, this invention constructs a fault event graph that integrates multimodal information and extracts multi-granular features from the fault event graph. It simultaneously obtains node-level fine-grained features that characterize local anomalies of nodes and graph-level coarse-grained features that characterize global fault modes of the system. This overcomes the problem of information one-sidedness caused by single-granularity feature extraction and provides more comprehensive feature support for root cause localization.
Smart Images

Figure CN122332173B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method, system and device for locating the root cause of microservice failures. Background Technology
[0002] Microservice architecture, with its high cohesion, low coupling, and scalability, has become the mainstream paradigm for building cloud-native applications. However, as systems grow in scale and the relationships between services become more complex, even minor anomalies in a single component can easily trigger cascading failures, leading to business interruptions. Therefore, quickly and accurately locating the root cause of failures in complex microservice topologies (i.e., root cause analysis) is crucial for ensuring system reliability.
[0003] Existing root cause localization techniques can be mainly divided into methods based on single-modal data and methods based on multi-modal data. Methods based on single-modal data, such as causal relationship analysis using indicator data, anomaly pattern mining using log data, or path analysis using call chain data, have achieved certain results in specific scenarios, but are limited by the one-sidedness of a single data source and are difficult to comprehensively capture complex faults caused by the interplay of multiple factors.
[0004] To overcome this limitation, the industry has begun to explore fusion solutions based on multimodal data. For example, Chinese patent application CN115357418A proposes a microservice fault detection method that generates anomaly event sequences by integrating multimodal data and uses a classification model for fault detection. However, this method primarily aims to determine whether a fault "exists," rather than to pinpoint "where the fault occurs." It employs a bag-of-words model for shallow feature extraction of the event sequences, making it difficult to capture the propagation path of faults and deep causal relationships within complex topologies.
[0005] Furthermore, Chinese patent application CN120909878A proposes a method for fault determination based on event causal graphs and agent-driven pre-trained models. This method improves the intelligence level of diagnosis to some extent by constructing event causal graphs and utilizing collaborative reasoning among multiple agents to locate root causes. However, this approach has the following limitations: First, its analysis of graph structures relies on the semantic understanding of large language models (LLMs) and the invocation of external tools. The feature extraction process is not an end-to-end learnable numerical process, leading to uncertainty in the reasoning results and limited efficiency. Second, this method reuses historical experience by retrieving external knowledge bases, which is a passive query and fails to achieve the active internalization and transfer of fault modes at the model parameter level. Third, it lacks a mechanism for explicit modeling and collaborative utilization of features of different granularities (such as local node anomalies and global system patterns).
[0006] In summary, existing technologies still have shortcomings in the following aspects: the one-sidedness of information representation (difficulty in deeply integrating multimodal data), the singularity of granular perspective (ignoring the synergy between global topological features and local anomalies), and the isolation of fault context (lacking effective knowledge transfer capabilities across fault samples). Overcoming these limitations and achieving accurate and robust root cause localization has become a pressing technical challenge in the current microservice operations and maintenance field. Summary of the Invention
[0007] To address the aforementioned technical problems, the technical solution adopted by this invention is as follows:
[0008] According to a first aspect of the present invention, a method for locating the root cause of microservice failures is provided, the method comprising the following steps:
[0009] Obtain multimodal monitoring data of the microservice system within the fault time window, and generate a multimodal feature vector corresponding to each microservice instance.
[0010] Based on the multimodal monitoring data, a topology graph is constructed with microservice instances as nodes and the call relationships between instances as edges. A corresponding multimodal feature vector is associated with each node in the topology graph to form a fault event graph.
[0011] Multi-granularity feature extraction is performed on the fault event graph to obtain node-level fine-grained features that characterize local anomalies of nodes and graph-level coarse-grained features that characterize global fault modes of the system.
[0012] Based on the graph-level coarse-grained features, the node-level fine-grained features are modulated to obtain enhanced node features that incorporate the global context.
[0013] Root cause localization is performed based on the enhanced node features, and root cause microservice instances are output.
[0014] According to a second aspect of the present invention, a microservice fault root cause localization system is provided, comprising:
[0015] The data acquisition module is used to acquire multimodal monitoring data of the microservice system within the fault time window.
[0016] The graph construction module is used to construct a topology graph based on the multimodal monitoring data, with microservice instances as nodes and the call relationships between instances as edges, and to associate each node in the topology graph with a corresponding multimodal feature vector to form a fault event graph.
[0017] The multi-granularity feature extraction module is used to extract multi-granularity features from the fault event graph to obtain node-level fine-grained features that characterize local anomalies of nodes and graph-level coarse-grained features that characterize global fault modes of the system.
[0018] The feature modulation module is used to modulate the node-level fine-grained features based on the graph-level coarse-grained features to obtain enhanced node features that incorporate the global context.
[0019] The location output module is used to perform root cause localization based on the enhanced node features and output the root cause microservice instance.
[0020] According to a third aspect of the present invention, an electronic device is provided, including a processor and a memory; the processor executes the steps of the method described in the first aspect of the present invention by invoking a program or instructions stored in the memory.
[0021] The present invention has at least the following beneficial effects:
[0022] First, this invention constructs a fault event graph that integrates multimodal information and extracts multi-granular features from the fault event graph. It simultaneously obtains node-level fine-grained features that characterize local anomalies of nodes and graph-level coarse-grained features that characterize global fault modes of the system. This overcomes the problem of information one-sidedness caused by single-granularity feature extraction and provides more comprehensive feature support for root cause localization.
[0023] Second, the present invention employs a feature modulation mechanism, which dynamically modulates node-level fine-grained features based on graph-level coarse-grained features. This allows the features of each node to retain their local anomaly information while integrating into the global fault context, thereby enhancing the separability of root cause nodes and non-root cause nodes in the feature space and improving the accuracy of root cause localization.
[0024] Third, this invention introduces a cross-graph attention mechanism during the model training phase. By interactively calculating the graph-level features of multiple fault event graphs within a batch, the model can learn the common patterns among different fault events and internalize these common patterns into model parameters, thereby improving the model's generalization ability for rare or complex faults. At the same time, the cross-graph attention mechanism is disabled during the inference phase to ensure the real-time nature of root cause localization.
[0025] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1A flowchart of a microservice fault root cause localization method provided in an embodiment of the present invention. Detailed Implementation
[0028] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0029] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of this invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0030] It should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the steps as sequential processes, many of these steps can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the steps can be rearranged. A process can be terminated when its operation is complete, but it may also have additional steps not included in the figures. A process can correspond to a method, function, procedure, subroutine, subroutine, etc.
[0031] Example 1: Method Example
[0032] To address the limitations of existing methods described in the background section, such as one-sided information representation, singular granular perspective, and isolated fault context, this invention provides a microservice fault root cause localization method based on multimodal data and multi-granular feature extraction. This method aims to construct a multi-layered, comprehensive root cause localization framework to overcome the shortcomings of existing technologies that only focus on local node states or single data views.
[0033] To achieve this objective, this invention first unifies the representation of heterogeneous monitoring data through a graph structure, mapping multimodal data such as metrics, logs, and call chains to the same topology to construct a fault event graph that integrates multi-source information. Second, it utilizes a hierarchical multi-granularity graph attention network (MG-GAT) to collaboratively extract fine-grained features at the node level and coarse-grained features at the graph level. The former characterizes the local abnormal states of microservice instances, while the latter depicts the global fault patterns of the entire system, thereby achieving hierarchical perception of fault propagation mechanisms. Finally, a cross-graph attention training strategy is introduced, enabling the model to discover common patterns among different fault events within a batch during training. This internalizes historical fault knowledge into model parameters, and the learned global contextual knowledge is transferred to the root cause localization task of the current fault through a feature modulation mechanism (FiLM), achieving context-aware localization from global patterns to local nodes.
[0034] Through the above-described technical means, this invention can achieve accurate and robust root cause localization in complex and dynamic microservice environments, effectively improving the accuracy of root cause identification and the model's generalization ability. The technical solution of this invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0035] like Figure 1 As shown in the figure, an embodiment of the present invention provides a microservice fault root cause localization method, which is implemented based on a root cause localization model. The root cause localization model includes a graph attention network, a feature modulation network, a cross-graph attention module, and a stacked multilayer perceptron. The cross-graph attention module is only enabled during the model training phase and disabled during the inference phase. The method includes the following steps:
[0036] S100: Obtain multimodal monitoring data of the microservice system within the fault time window, and generate multimodal feature vectors corresponding to each microservice instance.
[0037] This step acquires multimodal monitoring data of the microservice system within the failure time window. In this application, microservice refers to a software architecture style that divides a single application into a group of small, independently deployed service units. Each service unit runs in its own process, is built around business capabilities, and collaborates through lightweight communication mechanisms. Multiple microservices together constitute a microservice system.
[0038] Specifically, the microservice system refers to a distributed software architecture composed of multiple independently deployed, scalable, and collaborative microservice instances. Each microservice instance is responsible for implementing specific business functions, and the instances interact with each other through lightweight communication mechanisms (such as HTTP / RESTful API, RPC, etc.) to jointly complete complex business processing flows. A microservice instance is the smallest granularity service unit in the microservice architecture that can be independently deployed, run, and scaled. Each instance corresponds to a specific running process or container, has a unique instance identifier, carries specific business functions, and provides standardized service interfaces to the outside world.
[0039] The fault time window refers to a preset time period that is traced back or extended forward from the moment the fault is triggered, and is used to capture the system operating status related to the fault.
[0040] (a) Definition of multimodal monitoring data
[0041] The multimodal monitoring data includes at least two of the following types:
[0042] Metrics data: Time-series data reflecting the resource usage and performance of microservice instances, including but not limited to CPU utilization, memory utilization, disk I / O, network throughput, request latency, etc.
[0043] Log data: Semi-structured or unstructured text data that records the running status of microservice instances, business operations, and system events, including information such as timestamps, log levels, and message content;
[0044] Call chain data: Records the chain data of service call relationships and request processing, including fields such as Trace identifier, Span identifier, calling service, called service, call time, and response status.
[0045] (II) Anomaly Detection and Alarm Event Generation
[0046] Differentiated anomaly detection strategies are employed for different types of monitoring data to generate standardized alarm events:
[0047] Metric data anomaly detection: Anomaly detection is performed using the 3-sigma rule, generating metric alarm events that include instance identifier, metric name, and anomaly type.
[0048] Specifically, the mean and standard deviation of the indicator are calculated within a historical time window. When the current observation deviates from the mean by more than three times the standard deviation, it is considered an anomaly. The detected anomaly record is formatted as a triple of "Instance ID-Measure Name-Anomaly Type" as an alarm event for the indicator modality.
[0049] Log data anomaly detection: A dual strategy is adopted, combining explicit anomaly detection based on ERROR keywords and implicit anomaly detection based on low-frequency templates, to generate log alarm events containing instance identifiers and log keywords.
[0050] Specifically, the first layer is explicit anomaly detection, which identifies and clarifies error logs by matching predefined keywords such as "ERROR". The second layer is implicit anomaly detection, which calculates the average frequency and standard deviation of each log template based on its frequency distribution within a historical time window. When the frequency of a log template within the current time window is lower than the average frequency minus k times the standard deviation, the log template is determined to be abnormal, where k is a preset normal fluctuation range coefficient. The detected abnormal logs are structured into a "Instance ID-LogKey" tuple format as alarm events for the log modality. The LogKey is a unique identifier for the log template, obtained by parsing the original log message, and is used for classifying and indexing the structured log events.
[0051] Call chain data anomaly detection: The isolated forest algorithm is used for anomaly detection, generating call chain alarm events containing call instance identifiers.
[0052] Specifically, an isolated forest model is trained based on historical call chain data to learn the behavior patterns of normal call chains. For call chain data within the current time window, the trained model is used to identify abnormal features such as long-tail latency or error status codes. Detected abnormal call chain records are formatted as triplets of "call instance ID - called instance ID - interface" as alarm events for call chain modalities.
[0053] Alarm event standardization: Metric alarm events, log alarm events, and call chain alarm events are stored as standardized alarm events and associated using a unified event identifier (EventID). Each standardized alarm event includes at least an instance identifier, an event identifier, a modality type, and an exception information field.
[0054] (III) Feature Vector Embedding
[0055] Embedding technology is used to convert standardized alert events into dense vectors of fixed dimensions. Specifically, fastText embedding technology is used to map discrete text alert records into numerical feature vectors. Each microservice instance obtains a feature vector in each of the three modalities: metrics, logs, and call chains, forming the multimodal feature triple Hv=(H m v, H n v, H t v), where H m v, H n v, H tv corresponds to the feature embedding representations of the indicator modality, log modality, and call chain modality, respectively.
[0056] This completes the transformation from raw multimodal monitoring data to multimodal feature vectors for each microservice instance, providing input for the subsequent construction of the fault event graph.
[0057] S200 constructs a topology graph and associates multimodal features with nodes to form a fault event graph.
[0058] This step, based on the multimodal feature vector generated in step S100 and the call chain information in the multimodal monitoring data, constructs a topology graph with microservice instances as nodes and call relationships between instances as edges. It then associates each node in the topology graph with its corresponding multimodal feature vector, forming a fault event graph. Specifically, it includes the following sub-steps:
[0059] S201: Based on the call chain data in the multimodal monitoring data, the service dependency relationship within the fault time window is parsed out, and a topology graph structure with microservice instances as nodes and inter-instance call relationships as edges is generated.
[0060] Based on the call chain data within the fault time window, the call relationships between microservice instances and their attributes such as direction, frequency, and latency are analyzed. Each microservice instance is treated as a node in the topology graph, and the call relationships between instances are treated as edges between nodes, generating an initial topology graph structure G=<V, E>, where V is the set of microservice instance nodes and E is the set of call relationship edges.
[0061] S202, construct bidirectional edges in the topological graph structure.
[0062] To comprehensively capture the characteristics of fault propagation, bidirectional edges are constructed in the topology graph structure. These bidirectional edges include forward edges and reverse edges: forward edges are constructed along the direction of business calls to simulate the flow path of business call requests; reverse edges are constructed in the opposite direction of business calls to explicitly model the fault backtracking path, enabling subsequent root cause localization models to trace the fault propagation trajectory in reverse along the call chain.
[0063] S203, obtain multimodal feature vectors.
[0064] The multimodal feature vector corresponding to each microservice instance is obtained from step S100. The multimodal feature vector includes the indicator modality feature vector, the log modality feature vector, and the call chain modality feature vector. The feature vector of each modality is obtained through the anomaly detection and vectorization embedding process described in step S100, and a unified event identifier EventID has been assigned for cross-modal association.
[0065] S204, Align the multimodal feature vectors with the nodes of the topology graph and mount them to obtain the fault event graph.
[0066] Using the microservice instance identifier and event identifier EventID as the association key, the feature vectors of the same instance under different modalities are aligned, and the aligned feature vectors are then attached to the corresponding nodes of the topology graph constructed by S201-S202. After the above processing, each node v∈V in the fault event graph carries a multimodal feature triple.
[0067] At this point, the fault event graph G = <V, E, H>, which integrates multimodal information, is completed, providing a complete structured input for subsequent multi-granularity feature extraction and root cause localization.
[0068] Furthermore, in the process of constructing the fault event graph, this invention introduces a dynamic topology adaptation and incremental update mechanism to address scenarios such as dynamic changes in scale, instance scaling, and real-time changes in service topology faced by microservice systems in actual operation. In existing technologies, most solutions construct fault event graphs using static graph structures or full redrawing. In environments with high system dynamism and rapid update frequencies, this can easily lead to high computational overhead and graph construction lag, affecting the real-time performance and accuracy of root cause localization. The dynamic topology adaptation and incremental update mechanism of this invention specifically includes the following sub-steps:
[0069] S205 monitors the dynamic changes in the topology of the microservice system in real time, performs incremental updates to the subgraph of the changed parts, and obtains the features of the subgraph after the incremental update.
[0070] Real-time monitoring of the microservice system's operational status includes microservice instance startup and destruction, container scaling events, deployment version changes, and service call relationship adjustments. These operational status changes are collectively referred to as dynamic topology changes. When a dynamic topology change is detected, this step does not perform a full reconstruction of the currently used fault event graph. Instead, it identifies the changed service nodes and their directly associated edges, treating the subgraph formed by the changed node and its first-order neighbors as a local subgraph. Only this local subgraph undergoes feature recalculation, edge relationship updates, and node attribute iteration. By avoiding a full reconstruction of the original fault event graph, this step reduces the computational complexity of graph structure updates, minimizes resource consumption, and ensures that the updated fault event graph promptly reflects system topology changes.
[0071] S206, the incrementally updated subgraph features are fused with historical accumulated features to obtain a dynamically adaptive real-time fault event graph.
[0072] The incrementally updated subgraph features include the topology and operational status information of the changed node and its neighboring nodes at the current moment. At the same time, the system maintains the historical cumulative features of each node. These historical cumulative features record the evolution trend and abnormal patterns of each microservice instance over multiple time windows in the past, such as the long-term fluctuation patterns of indicators, the frequency distribution of log anomalies, and the statistical characteristics of call latency.
[0073] This step fuses the subgraph features with the historical accumulated features of the corresponding nodes, specifically using methods such as weighted fusion or cross-layer attention fusion. Through fusion, the features of each changed node in the updated fault event graph contain both real-time information at the current moment and retain the key patterns of that node during its historical operation. For nodes that have not changed, their historical accumulated features remain unchanged.
[0074] After the above processing, the fault event graph achieves dynamic adaptive updating: the graph structure reflects the changes in system topology in a timely manner, and the node features integrate real-time information and historical patterns, providing an accurate and stable graph structure foundation for subsequent root cause localization.
[0075] S300, perform multi-granularity feature extraction on the fault event graph.
[0076] This step performs multi-granularity feature extraction on the fault event graph constructed in step S200, obtaining node-level fine-grained features representing local node anomalies and graph-level coarse-grained features representing the global fault mode of the system. The fault event graph simultaneously contains coarse-grained structural information reflecting the global fault mode of the system and fine-grained information including the node's own state, modal attributes, and local topological relationships. To overcome the limitations of single-granularity feature extraction and achieve collaborative extraction of multi-granularity features, this step uses a multi-granularity graph attention network (MG-GAT) to process the fault event graph. The multi-granularity graph attention network includes node-level feature extraction branches and graph-level feature extraction branches, and the specific processing procedure is as follows.
[0077] S301, the node-level feature extraction branch performs node-level fine-grained feature extraction on the fault event graph.
[0078] Considering the significant semantic differences between metrics, logs, and call chain data, mixed encoding may lead to feature interference. Therefore, this step adopts an independent encoding strategy. For each modality in metrics, logs, and call chains, an independent graph attention network is used for feature learning. The core idea of the graph attention network is to assign different weights to different neighbors of a node through an attention mechanism, thereby focusing on neighbors more closely related to the current node when aggregating neighbor information. The specific processing procedure is as follows:
[0079] (1) For each modality, a learnable linear transformation is first performed on the original input features of each node in the fault event graph under that modality, mapping the original features to a new feature space to obtain the transformed feature representation. Taking the index modality as an example, a linear transformation is performed on the index features of each node; similarly, the same processing is performed on the log modality and the call chain modality. For node i∈V, its transformed feature representation h c i for:
[0080] h c i =Wc·h i .
[0081] Among them, h i Let Wc be the original input features of node i in the current modality, and h be the learnable weight matrix. c i Let i be the feature representation obtained after linear transformation of node i.
[0082] (2) The graph attention mechanism is used to calculate the unstandardized attention score between each node and its neighboring nodes.
[0083] For node i and its neighbor node j, the transformed features of both are concatenated, and then an unstandardized attention score is calculated using a learnable attention vector and a first non-linear activation function. This calculation process can be represented as: concatenating the transformed features of node i and the transformed features of its neighbor node j, performing a dot product with the learnable attention vector, and then processing the result through the first non-linear activation function to obtain the unstandardized attention score e. ij :e ij =LeakReLU(a T ·[h c i ||h c j ]), where e ij Let h be the unnormalized attention score between the i-th node and its j-th neighbor node in the fault event graph, j∈V(i), where V(i) is the set of neighbor nodes of the i-th node, and h is the unnormalized attention score between the i-th node and its j-th neighbor node. c i h represents the feature representation of the i-th node after linear transformation. c j Let a be the feature representation of the j-th neighbor node of the i-th node after linear transformation, and LeakReLU() be the first nonlinear activation function. T Let be a learnable attention vector, and || denote the vector concatenation operation.
[0084] (3) Use the Softmax function to normalize the attention scores of all neighbors of the node to obtain the standardized attention weights.
[0085] (4) The normalized attention weights are used as coefficients to perform a weighted summation of the features of the neighboring nodes after linear transformation, and then the updated features of each node are obtained through the second nonlinear activation function.
[0086] To enhance the model's expressive power, this step introduces a multi-head attention mechanism on top of the single-head attention mechanism: multiple independent attention heads are used to execute the aforementioned attention calculation process in parallel. Each attention head learns a different attention weight distribution, thereby capturing the relationships between nodes from different perspectives. Specifically, for each attention head, the aforementioned linear transformation, attention score calculation, attention weight normalization, and weighted summation processes are repeated. The features learned by multiple attention heads are concatenated along the channel dimension to serve as the node features output by the current layer.
[0087] .
[0088] X i Let be the feature representation of node i after processing by a single-layer graph attention network, which will be used as the input to the next layer of the graph attention network. σ is the second nonlinear activation function. The standardized attention weights between node i and its neighbor node j, calculated for the x-th attention head, are used to quantify the contribution of neighbor node j when updating the features of node i. This represents the learnable weight matrix corresponding to the x-th attention head in the multi-head attention mechanism, used to perform linear transformations on the features of neighboring nodes; x ranges from 1 to n, where n is the number of attention heads in the multi-head attention mechanism. The number of attention heads in the multi-head attention mechanism is set to 4 or 8, which can be adjusted according to the dimension of the graph-level features. In this embodiment, n is set to 4.
[0089] The first nonlinear activation function (LeakyReLU for calculating the attention score) and the second nonlinear activation function (σ after weighted summation) can be the same or different activation functions. To further enhance the node's ability to perceive multi-level neighbor features, this step uses iterative feature propagation through stacked multi-layer graph attention networks. Each layer uses the output of the previous layer as input, aggregating information in a wider neighborhood, enabling each node to fuse features from more distant neighbor nodes layer by layer. The number of stacked layers can be set according to the size and complexity of the fault event graph: 2 layers for small-scale graphs with fewer than 10 nodes; 3 layers for medium-scale graphs with 10-20 nodes; and 4 layers for large-scale graphs with more than 20 nodes. Experiments have verified that this layer setting scheme achieves a balance between computational efficiency and feature representation capability. After multi-layer stacking, each node finally obtains node-level fine-grained features that fuse its own state, modal attributes, and local topological information.
[0090] Through the above processing, the node-level fine-grained features of each microservice instance under the three modalities of metrics, logs, and call chains are obtained respectively.
[0091] S302, the graph-level feature extraction branch performs global max pooling on the node-level fine-grained features to obtain graph-level coarse-grained features.
[0092] After obtaining the node-level fine-grained features under each modality, in order to capture the system-level global failure mode, the node-level features need to be aggregated into graph-level features. Since failures in microservice systems are usually triggered by the extreme abnormal states of a few nodes, rather than the average deterioration of the state of all nodes, average pooling can easily smooth out critical abnormal signals. Therefore, this step adopts a global max pooling strategy to output the global maximum abnormal response feature representing the global failure mode, as the graph-level coarse-grained feature. Specifically, the maximum value is taken for each dimension of the node-level fine-grained feature, and the maximum response value in all dimensions is retained, thereby keenly capturing the most significant abnormal signals in the graph structure.
[0093] For each modality, global max pooling is performed on the node-level fine-grained features of all microservice instances under that modality to obtain graph-level coarse-grained features that characterize the failure modes of the entire system under that modality. This feature retains the global maximum response value in each dimension, highlighting the most significant feature information during the failure propagation process.
[0094] S303 concatenates multimodal features to obtain a unified multi-granularity feature representation.
[0095] Through the above processing, node-level features and graph-level features are obtained for three modalities: metrics, logs, and call chains. To achieve complementary fusion of multimodal information, the node-level feature vectors of each modality are concatenated along the channel dimension to generate a unified node-level feature representation; similarly, the graph-level feature vectors of each modality are concatenated along the channel dimension to generate a unified graph-level feature representation.
[0096] Thus, this step yields a unified feature representation with multi-granularity perspectives: node-level fine-grained features retain local anomaly information for each microservice instance across multiple modal dimensions, while graph-level coarse-grained features characterize the overall global failure mode of the system. This multi-granularity feature representation provides comprehensive and hierarchical feature support for subsequent root cause localization.
[0097] S400, based on the graph-level coarse-grained features, the node-level fine-grained features are modulated to obtain enhanced node features that incorporate the global context.
[0098] This step is implemented through the feature modulation network. This network uses the graph-level coarse-grained features extracted in step S300 as conditional information to dynamically modulate the node-level fine-grained features. This allows each node's features to retain its local anomaly information while integrating into the overall global fault context of the system, thereby enhancing the discriminative ability of node features and laying the foundation for effectively distinguishing root cause nodes from non-root cause nodes in subsequent root cause localization. Specifically, it includes the following sub-steps: S401, inputting the graph-level coarse-grained features into the feature modulation network to generate dynamic modulation parameters.
[0099] The graph-level coarse-grained feature f obtained in step S300 is input into the feature modulation network FiLM. Net Generate dynamic modulation parameters for each node i, including the scaling factor γ. i Translation factor β i The feature modulation network adopts a multilayer perceptron (MLP) architecture, and its specific structure and processing flow are as follows:
[0100] (a) Network input
[0101] The input to the feature modulation network is a graph-level coarse-grained feature f∈R. D Where D is the dimension of the graph-level feature. This feature integrates information from three modalities: metrics, logs, and call chains, representing the overall global failure mode of the system.
[0102] (II) Network Structure
[0103] The feature modulation network consists of stacked fully connected layers, specifically including:
[0104] Input layer: Receives graph-level coarse-grained features f, with D nodes.
[0105] Hidden layers: These consist of several fully connected layers, each followed by a non-linear activation function (such as ReLU). The number of neurons in a hidden layer can be set to D or adjusted empirically (e.g., 2D or D / 2), and the number of layers is typically set to 2 to 3. In this embodiment, the hidden layers of the feature modulation network are set to 2 layers, with the first layer having 2D neurons and the second layer having D neurons, and both layers using the ReLU activation function.
[0106] Output layer: Contains two parallel fully connected layers, one for generating the scaling factor vector and the other for generating the translation factor vector. The output layer has 2N neurons, where N is the number of nodes in the fault event graph. The output layer uses a linear activation function (i.e., no activation function) to ensure that the range of output parameters is unrestricted.
[0107] (III) Parameter Generation Process
[0108] For a graph-level coarse-grained feature f, the feature modulation network generates modulation parameters through the following calculations:
[0109] Forward propagation: f passes through the linear transformation and nonlinear activation of the input layer and hidden layer in sequence, extracting the modulation-related feature representation layer by layer.
[0110] Output separation: The output layer maps the final feature representation output by the hidden layer to two sets of parameters through a linear transformation: a scaling factor vector γ = [γ1, γ2, ..., γ]. i ,……,γ N Translation factor vector β=[β1,β2,……,β i ,……,β N ], where γ i and β i ∈R D1 D1 represents the dimension of the node-level fine-grained features, and R represents the set of real numbers.
[0111] Parameter allocation: The scaling factor vector and translation factor vector are split according to the node index to obtain the scaling factor and translation factor corresponding to each node, which are used for subsequent affine transformation of the node-level fine-grained features of that node.
[0112] (iv) Mathematical Expression
[0113] The computation process of the feature modulation network can be uniformly represented as: (γ,β)=MLP FiLM (f).
[0114] Among them, MLP FiLM () represents the mapping function of the multilayer perceptron mentioned above, whose parameters are learned end-to-end through the backpropagation algorithm during model training.
[0115] Through this network, coarse-grained graph features are converted into modulation parameter pairs corresponding to the number of nodes. Each modulation parameter pair is used to perform an affine transformation on the node-level fine-grained features of the corresponding node, thereby realizing the dynamic modulation of local node features by the global fault mode.
[0116] S402 uses dynamic modulation parameters to perform affine transformation on node-level fine-grained features to obtain enhanced node features.
[0117] Using the scaling and translation factors generated in step S401 for each node, perform an affine transformation on the node-level fine-grained features obtained in S300 to obtain enhanced node features:
[0118] Z i =γ i e i +β i .
[0119] in, Indicates element-wise multiplication, e i Z represents the node-level fine-grained feature of node i. i Enhanced node features for node i that incorporates the global context.
[0120] This affine transformation adjusts the amplitude of node features through a scaling factor and the bias of node features through a translation factor, so that each node feature retains its local anomaly information while being endowed with contextual information related to the global failure mode. After modulation, the separability of root cause nodes and non-root cause nodes in the feature space is enhanced, thus providing better feature input for subsequent root cause localization.
[0121] S500, perform root cause localization based on the enhanced node features, and output the root cause microservice instance.
[0122] This step is implemented using a stacked multi-layer perceptron. The stacked multi-layer perceptron performs root cause localization based on the enhanced node features obtained in step S400, identifying the root cause microservice instance causing the failure and its corresponding cause. The stacked multi-layer perceptron adopts a two-layer structure: the first layer calculates the anomaly score for each microservice instance and generates a root cause candidate list; the second layer predicts the failure cause type for each candidate instance, ultimately outputting a root cause localization result containing the root cause microservice instance and its corresponding failure cause.
[0123] Specifically, it includes the following sub-steps:
[0124] S501, the enhanced node features are input into a stacked multilayer perceptron to calculate the anomaly score for each microservice instance.
[0125] The enhanced node features Z of each node i obtained in step S402 are... i The input is classified using a stacked multilayer perceptron. The first layer of the multilayer perceptron enhances the node features Z. i As input, an anomaly score is calculated for each microservice instance, which quantifies the likelihood that the instance is the root cause. The higher the anomaly score, the greater the probability that the instance is the root cause of the failure.
[0126] S502, sort each microservice instance based on the anomaly score to generate a root cause candidate list.
[0127] Based on the anomaly score of each microservice instance calculated in step S501, all microservice instances are sorted in descending order to generate a root cause candidate list. Instances ranked higher in this root cause candidate list are more likely to be the root cause.
[0128] S503 inputs the root cause candidate list into the second layer multilayer perceptron, predicts the fault cause type of each candidate instance, and outputs the root cause localization result containing the root cause microservice instance and its corresponding fault cause.
[0129] The root cause candidate list generated in step S502 is fed into the second-layer multilayer perceptron. Based on the ranking information in the candidate list and the enhanced node features of each instance, the second-layer perceptron further predicts the most likely fault cause type for each candidate instance. Finally, the model outputs a root cause localization result containing the root cause microservice instance and its corresponding fault cause, providing operations personnel with accurate localization information to support timely remedial measures and minimize the impact of the fault.
[0130] Furthermore, after the S500 outputs the root cause microservice instance, this method also includes fault propagation path tracing and interpretable report generation steps to improve the interpretability of root cause localization results and operational decision support capabilities:
[0131] S510, based on the attention weights and fault propagation probabilities of a multi-granularity graph attention network, generates a fault causal propagation path by backtracking along the reverse edges of the fault event graph, starting from the root cause microservice instance.
[0132] This step utilizes the attention weights of each layer calculated by the multi-granularity graph attention network in step S300, and the fault propagation probability further calculated based on the attention weights, to perform reverse backtracking starting from the root cause microservice instance determined in step S500 and along the reverse edge constructed in step S202.
[0133] Specifically, for the current node, based on the attention weight distribution between it and its upstream neighbor nodes, the neighbor with the highest attention weight is selected as the upstream node for fault propagation; or the fault propagation probability is calculated based on the attention weight, and the propagation path with the highest probability is selected. The above process is repeated until the starting node of fault propagation (i.e., no upstream neighbor or propagation probability is lower than a preset threshold) is reached, thereby generating a complete fault causal propagation path, which records the complete trajectory of the fault propagating from the origin node to the root cause node.
[0134] The calculation of fault propagation probability based on attention weights can be implemented in the following ways:
[0135] In the multi-granularity graph attention network of step S300, attention weights between nodes are calculated for each graph attention layer. Taking the node-level feature extraction branch as an example, for node i and its neighbor node j, the normalized attention weight α... ij This represents the contribution of neighboring node j when updating the characteristics of node i. This weight reflects the strength of the influence of node j's state on node i.
[0136] Considering that the direction of fault propagation is usually opposite to the direction of business calls (i.e., the fault propagates from the callee to the caller), this invention utilizes the attention weights corresponding to the reverse edges to quantify the probability of fault propagation. Specifically, for a parent-child node pair connected by a reverse edge (let the child node be c and the parent node be q), during the node-level feature extraction process, the attention weight α of the parent node q on the child node c... qc The weight of c in q when aggregating neighbors can be interpreted as the strength of the propagation of the abnormal state of child node c to parent node q. Therefore, this attention weight can be directly used as the probability of fault propagation from child node c to parent node q, or it can be normalized and used as a probability value.
[0137] In actual backtracking, starting from the root cause microservice instance determined in step S500, all its upstream neighbor nodes (i.e., services that call this instance) are examined along the reverse edge. The attention weights corresponding to each upstream neighbor (i.e., the attention weights of the upstream node to the current node) are obtained, and these weights are normalized to obtain the fault propagation probability. The upstream node with the highest probability is selected as the previous hop for fault propagation, and this process is repeated until the propagation probability is lower than a preset threshold or a service without upstream neighbors is reached, thereby generating a complete fault causal propagation path.
[0138] S520, the fault causal propagation path, corresponding multimodal anomaly basis, and fault cause type are formatted in natural language to generate and output an interpretable root cause localization report.
[0139] This step integrates the fault causal propagation path generated in S510, the multimodal anomalies recorded at each node during S100-S200 (such as specific values of abnormal indicators, key information of log anomalies, and characteristics of call chain anomalies), and the fault cause type predicted in S503. It then formats this information using a preset natural language template to generate a structured, interpretable root cause localization report. The report must include at least the following:
[0140] Root causes of failures in microservice instances and the types of failure causes;
[0141] The complete path of a fault propagating from the origin node to the root node;
[0142] Key anomaly evidence at each node along the path (such as "CPU utilization suddenly increased to 95%" or "the frequency of database connection failure error logs surged").
[0143] Recommended repair measures are suggested.
[0144] The report is output in natural language that is readable by operations and maintenance personnel, providing intuitive and comprehensive decision support for troubleshooting and system recovery.
[0145] Furthermore, the method of the present invention includes a training phase and an inference phase. The training phase implements a cross-graph attention mechanism to enhance the model's generalization ability, while the inference phase disables this mechanism to ensure real-time performance.
[0146] The training phase of the root cause localization model includes the following steps:
[0147] S600 applies a cross-graph attention mechanism to the coarse-grained graph-level features of multiple fault event graphs within a batch. By calculating and aggregating the similarities between different graphs, it obtains enhanced graph-level features that include contextual information of other fault events within the batch.
[0148] To overcome the limitations of the traditional "single-instance analysis" paradigm and enable the model to capture and utilize common knowledge among different failure events, a cross-graph attention mechanism is introduced during the model training phase, implemented through the aforementioned cross-graph attention module. This module enhances the context of graph-level representations by mining the correlations between graphs of different failure events within a batch, guiding the model to learn more robust and discriminative graph-level features. Specifically, it includes the following sub-steps:
[0149] S601 performs linear projection on the graph-level coarse-grained features of each fault event graph within the batch to generate query vector, key vector, and value vector.
[0150] During training, given a batch containing M fault event graph samples, its corresponding graph-level coarse-grained feature set is {f1, f2, ..., f...}. M For each graph-level coarse-grained feature fr Perform linear projection on (r=1……M) to generate query vectors Q respectively. r Key vector K r Sum vector V r Q r =f r ·W Q ;K r =f r ·W K V r =f r ·W V .
[0151] Among them, W Q W K W V These are learnable weight matrices, whose dimensions are pre-defined based on the dimensions of the graph-level coarse-grained features.
[0152] S602, calculate the similarity between the query vector of the current fault event graph and the key vectors of other graphs in the batch, and perform a weighted summation of the value vectors of other graphs based on the similarity to obtain the enhanced graph-level features of the current graph.
[0153] First, calculate the query vector Q of the current fault event graph. r The key vector K of the batch and all fault event graphs (including itself) p The similarity between (p=1……M) is determined by a scaling factor (d). k ) 1 / 2 (d) k The attention score α is obtained by scaling the dimension of the key vector and then normalizing it using the Softmax function. rp :
[0154] α rp = .
[0155] This score quantifies the correlation between fault event graph r and fault event graph p in terms of global fault modes, where, The key vector K represents the graph of the p-th fault event. p The transpose of .
[0156] Subsequently, based on the aforementioned attention scores, the value vectors of all fault event graphs are weighted and summed to obtain the enhanced graph-level features of the current fault event graph. : V p This represents the value vector of the p-th fault event graph.
[0157] S603, the enhanced graph-level features are fused with the original coarse-grained graph-level features through residual connections to obtain the updated graph-level features f.r updated :f r updated =f r + .
[0158] By performing a residual connection between the enhanced graph-level features and the original coarse-grained graph-level features, updated graph-level features that incorporate global context are obtained. This residual connection allows the model to further incorporate contextual information from other fault events within the batch, while retaining the original graph-level features.
[0159] S700, the enhanced graph-level features are used to replace or combine the original graph-level coarse-grained features to modulate the node-level fine-grained features.
[0160] During the training phase of the root cause localization model, the updated graph-level features obtained in step S603 replace the original coarse-grained graph-level features, and these are input into the feature modulation network to modulate the fine-grained node-level features. The specific modulation process is the same as in step S400: the updated graph-level features are input into the feature modulation network to generate dynamic modulation parameters for each node, and affine transformations are performed on the fine-grained node-level features to obtain enhanced node features. In this way, the root cause localization model internalizes the common knowledge of other fault events within the batch into modulation parameters during training, thereby improving the feature discrimination capability.
[0161] During the training of the root cause localization model, the cross-entropy loss function is used as the optimization objective, combined with the Adam optimizer for end-to-end parameter learning. The initial learning rate is set to 0.001, and it decays to 0.1 every 10 training epochs. Through this setting, the model internalizes the common knowledge of other fault events within the batch into the network parameters during training, thereby improving the discriminative ability of the features.
[0162] During the inference phase of the root cause localization model, the aforementioned cross-graph attention mechanism is disabled, and the graph-level coarse-grained features of the current fault event graph are directly used for modulation to meet real-time requirements. The common fault patterns learned by the model through the cross-graph attention mechanism during the training phase are already encoded in the model parameters, so efficient and accurate root cause localization can be achieved without additional computation during the inference phase.
[0163] Example 2: Experimental Verification
[0164] To verify the effectiveness of the microservice fault root cause localization method based on multimodal data and multigranular feature extraction (hereinafter referred to as MMG-RCL) described in this invention, this embodiment conducted experiments on three publicly available microservice system datasets and compared them with existing baseline methods.
[0165] 1. Experimental Setup
[0166] Three open-source microservice datasets were used: D1 (GAIA dataset, 10 microservices, 1099 failure events), D2 (SockShop system, 14 microservices, 654 failure events), and D3 (HotelReservation system, 18 microservices, 490 failure events). All failures were injected by engineers through methods such as CPU stress, network packet loss, and database connection exhaustion, covering a variety of typical failure types.
[0167] The comparison methods include: iSQUAD based on logs, TraceRCA based on call chains, and DiagFusion, Eadro, and TVDiag based on multimodal approaches. AC@k (the probability that the first k results contain the root cause) and MRR@k (mean reciprocal rank) are used as evaluation metrics.
[0168] 2. Overall Performance Comparison
[0169] Experimental results show that MMG-RCL achieves optimal performance on all datasets. Taking the AC@1 metric as an example, it achieves scores of 0.81, 0.85, and 0.88 on D1, D2, and D3, respectively, representing a 2%-6% improvement over the best baseline TVDiag. Compared to single-modal methods, the improvement exceeds 60%, fully validating the effectiveness of multimodal data fusion. Among multimodal methods, DiagFusion and Eadro, which employ early fusion strategies, show significantly lower performance than our method, indicating that the strategy of independent encoding and late fusion can better preserve the fine-grained information of each modality.
[0170] 3. Efficiency Analysis
[0171] In CPU-based testing, the average inference time per sample is between 12 and 18 ms. Since cross-graph attention is only used during the training phase, no additional computation is required during the inference phase, which meets the real-time root cause localization requirements of microservice systems.
[0172] 4. Conclusion
[0173] The experimental results above demonstrate that this invention significantly improves the accuracy and robustness of microservice failure root cause localization through multimodal data fusion, multi-granularity feature collaborative extraction, and cross-graph attention enhancement, outperforming existing methods on multiple datasets.
[0174] Example 3: Application Scenario Example
[0175] This invention can be widely applied to various business systems employing microservice architecture, including but not limited to civil aviation passenger service systems, e-commerce transaction systems, online payment platforms, telecommunications business support systems, industrial internet platforms, and core financial transaction systems—scenarios requiring rapid microservice fault location. These systems typically contain dozens to hundreds of microservice instances, with complex inter-service call relationships. An anomaly in a single service can easily trigger cascading failures, severely impacting business continuity and user experience. The following detailed explanation uses a civil aviation passenger service system as an example to illustrate the specific application of this invention.
[0176] I. Overview of Application Scenarios
[0177] This embodiment is applied to the passenger service system of a large airline. The system is deployed using a microservice architecture, comprising 32 microservice instances, including flight search, fare calculation, order management, payment processing, check-in service, and membership management, deployed on a Kubernetes container cloud platform. These microservice instances interact via RESTful APIs to jointly complete the entire business process for passengers, from flight search to ticket purchase. The system handles over 5 million passenger requests daily, placing extremely high demands on system stability and fault recovery speed.
[0178] II. Fault Scenario Setting
[0179] During a peak business period on a certain day, the system suddenly malfunctioned: some passengers reported being unable to complete their ticket payments, with payment requests timing out or returning system errors. The operations and maintenance monitoring system triggered an alarm, requiring rapid identification of the root cause of the fault within the complex service call chain.
[0180] III. Multimodal Monitoring Data Acquisition
[0181] Within the fault time window (5 minutes before and 5 minutes after the fault trigger time), the following multimodal monitoring data were collected:
[0182] Metrics data: Collect metrics such as CPU utilization, memory usage, JVM GC count, HTTP request latency, and error rate for each microservice instance, with a sampling frequency of 30 seconds per instance.
[0183] Log data: Collect business logs and system logs from each microservice instance, including log information at different levels such as INFO, WARN, and ERROR.
[0184] Call chain data: Collect complete call chain information within the fault period, and record the microservice instances that each request passes through, the call time, the response status code, etc.
[0185] IV. Fault Event Diagram Construction
[0186] Based on the collected call chain data, service dependencies are parsed, and a topology graph structure is constructed with 32 microservice instances as nodes and the call relationships between instances as edges. Bidirectional edges are built in the topology graph: forward edges simulate the direction of business calls (e.g., "order service → payment service"), and reverse edges are used for fault backtracking path modeling.
[0187] Anomaly detection was performed on three modalities of data:
[0188] Anomaly detection: Using the 3-sigma rule, we detected a sudden increase in CPU utilization of the "Payment Service" instance to 85% (historical average is 30%), and a sudden increase in JVM GC frequency from the normal value of 5 times / minute to 120 times / minute.
[0189] Log anomaly detection: ERROR keyword detection revealed a large number of "database connection pool exhausted" errors in the "Payment Service" log; low-frequency template detection revealed that the frequency of the "Payment Request Timeout" log template in the "Order Service" increased from the normal value of 0 times / minute to 300 times / minute.
[0190] Call chain anomaly detection: Using the isolated forest algorithm, it was found that in the call chain involving "payment service", 95% of the request response time exceeded 5 seconds (the normal value is 200ms), and the returned status codes were mostly 500 or 504.
[0191] The above anomaly detection results are generated into standardized alarm events, and a unique event identifier (EventID) is assigned to each alarm event. Alarm events of the same instance in different modalities are associated using EventIDs to form a multimodal feature vector for each microservice instance.
[0192] V. Multi-granularity feature extraction
[0193] The constructed fault event graph is input into the Multi-Granularity Graph Attention Network (MG-GAT) for multi-granularity feature extraction:
[0194] Node-level feature extraction: For each microservice instance, a stacked 3-layer graph attention network is used to aggregate neighbor node information along the call relationship edges, resulting in fine-grained node-level features that integrate the microservice's own state and local topology information. Taking the "Payment Service" as an example, its node features integrate its own CPU anomaly indicators, database connection pool error logs, and timeout log information from its upstream "Order Service".
[0195] Graph-level feature extraction: Global max pooling is performed on the node-level features of all 32 nodes, retaining the maximum response value for each dimension to obtain graph-level coarse-grained features characterizing the global failure modes of the system. This feature highlights the most significant signals of anomalies in the entire system—in this example, the maximum response values are concentrated on the dimensions related to "payment services" and its database.
[0196] VI. Feature Modulation and Root Cause Localization
[0197] The graph-level coarse-grained features are input into a feature modulation network (2-layer MLP, 128 hidden layer dimensions) to generate dynamic modulation parameters (scaling and translation factors) for each node. These parameters are then used to perform affine transformations on the node-level fine-grained features to obtain enhanced node features that incorporate the global context.
[0198] The enhanced node features are input into a stacked multilayer perceptron (the first layer calculates the anomaly score, and the second layer predicts the cause of the failure) to obtain the root cause localization results. Analysis shows that the payment service ranks first in the root cause candidate list, with an anomaly score of 0.96, and the predicted cause of failure is database connection pool exhaustion; the order service ranks second, with an anomaly score of 0.78, and the predicted cause of failure is timeout in a dependent service; the membership service ranks third, with an anomaly score of 0.45, and its status is judged as normal fluctuation. The anomaly scores of the remaining microservice instances decrease sequentially and are not listed individually. The localization results show that the payment service failed due to database connection pool exhaustion, while the timeout in the order service is a cascading effect rather than the root cause.
[0199] The location result indicated that the "payment service" malfunctioned due to "database connection pool exhaustion," and the timeout of its upstream "order service" was a cascading effect rather than the root cause.
[0200] VII. Verification of Technical Effects
[0201] Practical verification shows that the positioning results of this invention are consistent with those obtained through manual investigation. Compared with traditional methods:
[0202] Using a single indicator for analysis (focusing only on sudden CPU spikes) can lead to misjudgments of "insufficient computing resources," resulting in incorrect capacity expansion.
[0203] Using a single log analysis method (focusing only on the ERROR log) will pinpoint "database connection failed", but it cannot determine which service caused the failure.
[0204] Using this method, it only takes 15ms from fault triggering to root cause output, and accurately identifies the root cause of "Payment Service - Database Connection Pool Exhaustion". Based on this, the operation and maintenance personnel can adjust the database connection pool configuration in a timely manner and restore the service within 5 minutes.
[0205] This embodiment demonstrates that the present invention can effectively address cascading failure scenarios in complex civil aviation passenger service systems, quickly and accurately locating the root cause microservice instance. Similarly, in other microservice application scenarios such as e-commerce, financial payments, and telecommunications operations, the present invention can also leverage its technical advantages to provide strong support for ensuring the continuity of core businesses.
[0206] Example 4: System Example
[0207] This embodiment provides a microservice fault root cause localization system, including:
[0208] The data acquisition module is used to acquire multimodal monitoring data of the microservice system within the fault time window. The microservice system consists of multiple independently deployed and collaborative microservice instances; the fault time window is a preset time period that traces backward or extends forward from the fault triggering time; the multimodal monitoring data includes at least two of the following: metric data, log data, and call chain data.
[0209] The graph construction module is used to construct a topology graph based on the multimodal monitoring data, with microservice instances as nodes and call relationships between instances as edges. It also associates a corresponding multimodal feature vector with each node in the topology graph to form a fault event graph. Specifically, the graph construction module generates a topology graph structure based on service dependency relationships parsed from the call chain data, constructs bidirectional edges containing forward and reverse edges, and uses a unified event identifier to align feature vectors of different modalities and attach them to the corresponding nodes to form the fault event graph.
[0210] A multi-granularity feature extraction module is used to extract multi-granularity features from the fault event graph to obtain node-level fine-grained features representing local anomalies and graph-level coarse-grained features representing global fault modes of the system. The multi-granularity feature extraction module includes a multi-granularity graph attention network, which further includes:
[0211] The node-level feature extraction branch aggregates neighbor node information along the edges of the fault event graph using stacked graph attention layers, outputting the node-level fine-grained features. This branch encodes each modality—metrics, logs, and call chains—using independent graph attention networks, and concatenates the multimodal encoding results as the node-level fine-grained features.
[0212] The graph-level feature extraction branch is used to perform global max pooling on the node-level fine-grained features, taking the maximum value for each dimension of the node-level fine-grained features, retaining the maximum response value in all dimensions, and using the vector composed of the maximum response values as the graph-level coarse-grained features representing the global fault mode.
[0213] The feature modulation module is used to modulate the node-level fine-grained features based on the graph-level coarse-grained features, obtaining enhanced node features that incorporate the global context. The feature modulation module employs a feature modulation mechanism, inputting the graph-level coarse-grained features into the feature modulation network to generate dynamic modulation parameters for each node, including scaling and translation factors. It then uses these scaling and translation factors to perform an affine transformation on the corresponding node's node-level fine-grained features to obtain the enhanced node features.
[0214] The location output module is used to perform root cause localization based on the enhanced node features and output the root cause microservice instance. The location output module inputs the enhanced node features into a stacked multilayer perceptron, calculates the anomaly score for each microservice instance, sorts the instances based on the anomaly scores to generate a root cause candidate list, predicts the fault cause type for each candidate instance, and finally outputs the root cause localization result containing the root cause microservice instance and its corresponding fault cause.
[0215] Furthermore, the system also includes a cross-graph attention enhancement module during the model training phase. This module applies a cross-graph attention mechanism to the coarse-grained graph features of multiple fault event graphs within a batch: it performs linear projection on each coarse-grained graph feature to generate a query vector, key vector, and value vector; it calculates the similarity between the query vector of the current graph and the key vectors of other graphs in the batch, and performs a weighted summation of the value vectors of other graphs based on this similarity to obtain the enhanced graph-level features of the current graph; it then fuses the enhanced graph-level features with the original coarse-grained graph features through residual connections to obtain updated graph-level features, which are used to replace or combine with the original coarse-grained graph features for feature modulation. During the model inference phase, the cross-graph attention enhancement module is deactivated, and the current fault event graph's own coarse-grained graph features are used directly for modulation to meet real-time requirements.
[0216] The specific implementation methods and data processing flows of each of the above modules are described in the corresponding steps of the method embodiments of the present invention, and will not be repeated here.
[0217] Example 5: Electronic Device Example
[0218] This invention also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being configured to perform the method described in this invention.
[0219] This invention also provides a computer-readable storage medium storing computer-executable instructions for performing the methods described in this invention.
[0220] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this invention can be achieved, and this is not limited herein.
[0221] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method for locating the root cause of microservice failures, characterized in that, The method includes the following steps: Acquire multimodal monitoring data of the microservice system within the fault time window, and generate a multimodal feature vector corresponding to each microservice instance; the multimodal monitoring data includes at least two of the following types: indicator data, log data, and call chain data; Based on the multimodal monitoring data, a fault topology graph is constructed with microservice instances as nodes and the call relationships between instances as edges. A corresponding multimodal feature vector is associated with each node in the topology graph to form a fault event graph. Multi-granularity feature extraction is performed on the fault event graph to obtain node-level fine-grained features that characterize local anomalies of nodes and graph-level coarse-grained features that characterize global fault modes of the system. Based on the graph-level coarse-grained features, the node-level fine-grained features are modulated to obtain enhanced node features that incorporate the global context. Root cause localization is performed based on the enhanced node features, and root cause microservice instances are output. The fault event diagram is obtained through the following steps: Based on the call chain data in the multimodal monitoring data, the service dependency relationship within the fault time window is parsed out, and a topology graph structure with microservice instances as nodes and inter-instance call relationships as edges is generated; In the topology graph structure, bidirectional edges are constructed, including forward edges that simulate normal business flow and reverse edges that are used to model fault backtracking paths; The feature vectors of different modalities are aligned using a unified event identifier and attached to the corresponding nodes of the topology graph to obtain the fault event graph.
2. The method according to claim 1, characterized in that, The multi-granularity feature extraction of the fault event graph includes: The fault event graph is processed using a multi-granularity graph attention network; The multi-granularity graph attention network includes node-level feature extraction branches and graph-level feature extraction branches; The node-level feature extraction branch is used to aggregate neighbor node information along the edges of the fault event graph through stacked graph attention layers, and output the node-level fine-grained features. The graph-level feature extraction branch is used to perform global max pooling on the node-level fine-grained features and output the global maximum anomaly response feature that characterizes the global fault mode, which serves as the graph-level coarse-grained feature.
3. The method according to claim 2, characterized in that, The node-level feature extraction branch contains multiple parallel sub-networks, each corresponding to a modality. Each sub-network is used to encode the features of a node under that modality using an independent graph attention mechanism, and the features encoded by each modality are concatenated to obtain the node-level fine-grained features.
4. The method according to claim 1, characterized in that, During the training phase, the method further includes: A cross-graph attention mechanism is applied to the coarse-grained graph features of multiple fault event graphs within a batch. By calculating and aggregating the similarity between different graphs, an enhanced graph-level feature containing contextual information of other fault events within the batch is obtained. The enhanced graph-level features are used to replace or combine with the original graph-level coarse-grained features to modulate the node-level fine-grained features.
5. The method according to claim 4, characterized in that, During the inference phase, the cross-graph attention mechanism is disabled, and the node-level fine-grained features are directly modulated using the graph-level coarse-grained features extracted from the current fault event graph.
6. The method according to claim 1 or 4, characterized in that, Modulating the node-level fine-grained features based on the graph-level coarse-grained features specifically involves: The graph-level coarse-grained features are input into a feature modulation network to generate dynamic modulation parameters for each node, including scaling factors and translation factors. The enhanced node features are obtained by performing an affine transformation on the node-level fine-grained features of the corresponding node using the scaling factor and translation factor.
7. The method according to claim 1, characterized in that, Based on the enhanced node features, root cause localization is performed, and root cause microservice instances are output, including: The enhanced node features are input into a stacked multilayer perceptron to calculate the anomaly score for each microservice instance; Based on the anomaly scores, each microservice instance is sorted to generate a root cause candidate list; The output includes the root cause microservice instance and the root cause of the failure corresponding to the root cause service instance.
8. A microservice fault root cause localization system, characterized in that, include: The data acquisition module is used to acquire multimodal monitoring data of the microservice system within the fault time window; The multimodal monitoring data includes at least two of the following types: indicator data, log data, and call chain data; The graph construction module is used to construct a topology graph based on the multimodal monitoring data, with microservice instances as nodes and the call relationships between instances as edges, and to associate each node in the topology graph with a corresponding multimodal feature vector to form a fault event graph. The multi-granularity feature extraction module is used to extract multi-granularity features from the fault event graph to obtain node-level fine-grained features that characterize local anomalies of nodes and graph-level coarse-grained features that characterize global fault modes of the system. The feature modulation module is used to modulate the node-level fine-grained features based on the graph-level coarse-grained features to obtain enhanced node features that incorporate the global context. The location output module is used to perform root cause localization based on the enhanced node features and output the root cause microservice instance. The fault event diagram is obtained through the following steps: Based on the call chain data in the multimodal monitoring data, the service dependency relationship within the fault time window is parsed out, and a topology graph structure with microservice instances as nodes and inter-instance call relationships as edges is generated; In the topology graph structure, bidirectional edges are constructed, including forward edges that simulate normal business flow and reverse edges that are used to model fault backtracking paths; The feature vectors of different modalities are aligned using a unified event identifier and attached to the corresponding nodes of the topology graph to obtain the fault event graph.
9. An electronic device, characterized in that, Including processor and memory; The processor executes the steps of the method as described in any one of claims 1 to 7 by invoking programs or instructions stored in the memory.
Citation Information
Patent Citations
Micro-service fault detection method and device, storage medium and computer equipment
CN115357418A
Fault determination method and device of micro-service system, product and electronic equipment
CN120909878A
Micro-service system fault positioning method and system based on multi-modal data fusion
CN121706001A
Multimodal micro-service system fault diagnosis method based on heterogeneous graph multi-stage fusion
CN122046197A