A Fault Location Method for Microservice Applications by Learning from Link Data
By adopting the Transformer-BiLSTM model with spatiotemporal characteristics in the microservice system, combined with distributed tracking and multimodal adaptive gating mechanism, the problem of complexity of fault detection and positioning of microservice systems is solved, efficient fault prediction and positioning is achieved, and the stability and reliability of the system are improved.
Patent Information
- Application Number
- CN202510377832.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-28
AI Technical Summary
The fault detection and positioning of microservice systems are complex, and the existing technology is difficult to effectively capture the spatio-temporal characteristics of data, making it difficult to detect and position the fault detection and positioning, and the abnormality of a single service node is difficult to detect in time.
The Transformer-BiLSTM model integrating spatiotemporal features is adopted to predict and locate the microservice system. Through hybrid fault injection strategies, distributed tracking of Kubernetes' Istio service mesh, joint spatio-temporal attenuation factor, time-aware multi-head attention mechanism and multimodal adaptive gating mechanism, a model that can learn the spatio-temporal characteristics of microservice systems is constructed.
It significantly improves the accuracy and efficiency of fault detection and positioning of microservice systems, and can dynamically adjust the number of service instances to ensure system stability and reliability.
Smart Images

Figure CN119883714B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of distributed architecture fault location, and particularly relates to a method for fault location of microservice applications by learning from link data. Background Art
[0002] With the widespread adoption of distributed systems and microservice architectures in modern enterprises, the complexity of service deployment, maintenance, and expansion has reached an unprecedented level. Due to its excellent scalability, flexibility, and modularity, the microservice architecture has become the core technology for supporting large-scale distributed applications. However, the inherent dynamics, distributed characteristics, and heterogeneity of microservice systems make fault detection and location particularly complex. In a microservice environment, the tight dependencies between services and the high requirements for real-time response mean that any system fault can quickly spread, leading to a sharp decline in performance or even a complete system collapse. Therefore, developing a technology that can efficiently and accurately detect and locate microservice faults has become a key challenge in ensuring system stability and reliability.
[0003] Traditional microservice fault detection and location methods mainly rely on machine learning and deep learning technologies, including support vector machines (SVMs), random forests, long short-term memory networks (LSTMs), and CNN-LSTMs. These technologies usually identify anomalies based on single log data or by monitoring the usage of resources such as CPU, memory, and network. However, when dealing with microservice call relationships, these methods do not fully analyze the service topology graph in depth. Although LSTMs and CNN-LSTMs can handle time series data, and CNN-LSTMs can also handle spatial data, they still have deficiencies in capturing the spatio-temporal characteristics of data. In addition, with the continuous expansion of the microservice architecture, it is not only difficult to detect anomalies in individual service nodes in a timely manner, but once they occur, they often trigger a chain reaction, greatly increasing the overall system fault diagnosis difficulty. Summary of the Invention
[0004] In view of the above problems, the present invention proposes a method for fault location of microservice applications based on link data learning. This method uses a Transformer-BiLSTM model that fuses spatio-temporal features to predict and locate faults in microservice systems. To meet the requirements of actual application scenarios, this model is integrated into the microservice system to achieve real-time fault prediction and location functions for online applications.
[0005] The method of the present invention is mainly reflected in two aspects of microservice fault prediction and location. On the one hand, a hybrid fault injection strategy is used to simulate real application scenarios, and the Istio service mesh of Kubernetes is used to collect microservice link data through its distributed tracing and preprocess the data. On the other hand, by learning microservice application fault characteristics from the link data, a Transformer-BiLSTM model integrating spatio-temporal features is used to predict potential faults and locate faults in microservice applications.
[0006] S1. To simulate real usage scenarios, the present invention performs hybrid fault injection on an open-source microservice system, and then uses the Istio service mesh of Kubernetes to collect microservice link data through its distributed tracing and preprocess the data.
[0007] Further: The step 1 includes the following sub-steps:
[0008] S1-1. Perform hybrid fault injection on the open-source microservice system.
[0009] S1-2. Use the Istio service mesh of Kubernetes to collect microservice link data through its distributed tracing and perform data processing.
[0010] S1-3. Data preprocessing:
[0011] First, define the link data matrix as , where T is the time step, N is the number of microservices, and M is the feature dimension.
[0012] Next, construct a call relationship graph between services and set a static service call adjacency matrix . If service i calls service j, then , otherwise it is 0.
[0013] S2. Construction of a fault prediction and location model.
[0014] After data preprocessing, the present invention performs fault location of microservice applications through a Transformer-BiLSTM model integrating spatio-temporal features. This model combines spatio-temporal features and can fully learn the time-dependent relationship and spatial structure features in the microservice system.
[0015] S2-1. Perform temporal embedding and causal position encoding on the preprocessed data.
[0016] In a microservice system, the propagation of faults has obvious causal temporal characteristics. To accurately capture and understand this causal relationship, the causal position can be encoded, and the specific formula is as follows:
[0017]
[0018] Among them, learns the causal weight through historical anomaly data, indicates that if the time at time is the potential cause, it is set to 1, otherwise 0. is the standard sine position encoding vector.
[0019] In addition, the dynamic causal injection technology is also used to simulate and observe the responses of the microservice system in the face of configuration changes, traffic fluctuations, changes in inter-service calls, etc. Through this technology, the dynamic changes of the system can be observed in real time, and the causal relationship model of the present invention can be further verified and adjusted. The specific formula is as follows:
[0020]
[0021] Among them, is the input feature matrix, is the embedding layer weight matrix.
[0022] S2-2, Time-Aware Multi-Head Attention Mechanism.
[0023] A time-aware multi-head attention mechanism is introduced, which enhances the spatio-temporal perception ability of the model by integrating the call path distance of the service into the attention calculation.
[0024] First, by combining the service space topology relationship graph and the time decay factor, a spatio-temporal joint decay factor is constructed. This decay factor not only considers the time factor but also incorporates the spatial distance between services, thus realizing the joint perception of spatio-temporal information. The specific formula is as follows:
[0025]
[0026] Among them, is the shortest path distance from service i to j, , and are learnable parameters, and represent the timestamps of service i and service j.
[0027] Next, the attention mechanism is optimized using the above spatio-temporal joint decay factor. The specific formula is as follows:
[0028]
[0029] Among them, Q, K, and V are the query, key, and value respectively.
[0030] Finally, this optimized attention mechanism is extended to a multi - head setting to enhance the model's expressive power and generalization ability. Each head processes different representation sub - spaces, thereby capturing multi - faceted features of the input data. The final output of multi - head attention is the concatenation of the outputs of each head.
[0031] S2 - 3. Perform BiLSTM time - series modeling based on the output of the time - aware multi - head attention mechanism to obtain the final time features.
[0032] Process the link data of microservices through Transformer, and then input the high - dimensional features that can reflect the dynamics of time series into the BiLSTM layer. The time - series features output by Transformer are input into the BiLSTM network to capture local time dependencies. BiLSTM performs time - series modeling through forward and backward paths, and finally obtains the final time features by concatenating the bidirectional hidden states. 。
[0033] S2 - 4. Combine the spatio - temporal hyper - graph attention mechanism to extract spatial features.
[0034] When extracting spatial features, first construct node features and map the time features into the graph space. This process involves calculating the attention scores of each node for its neighbors in order to more accurately capture the interactions and information flows between nodes. The specific formula is as follows:
[0035]
[0036] Next, adopt the spatio - temporal hyper - graph attention mechanism to reflect the collaborative anomalies of multiple services. Each service and all its upstream dependent services form a hyper - edge is the set of direct upstream services of service i, is the hyper - edge of service i. The specific formula is as follows:
[0037]
[0038] Among them, is the max - pooling operation, is the activation function, is the attention coefficient vector.
[0039] Finally, to solve the complex relationships of services in the microservice architecture, aggregate the hyper - edge features to express the dependency relationships, thereby helping to detect service anomaly problems. The specific formula is as follows:
[0040]
[0041] S2-5. Based on spatial features, perform multimodal adaptive fusion through a multimodal adaptive gating mechanism.
[0042] When dealing with a large amount of data generated by inter-service communication, such as log files, performance metrics, and network traffic, etc., a multimodal adaptive gating mechanism is proposed to effectively analyze this data. A triple gating system is designed to optimize the fusion weights for temporal features, spatial features, and cross features.
[0043] First, to improve the accuracy of the model, project the cross-attention features and calculate a comprehensive cross-attention mechanism. This mechanism involves the application of an attention weight matrix, and the specific formula is as follows:
[0044]
[0045] Among them, 、 and are the attention weight matrices, is the temporal feature, is the spatial feature, is the cross feature, is the feature dimension.
[0046] Next, introduce the multimodal gating vector. This vector helps the model adjust the weight of information according to the characteristics of different data, and the specific formula is as follows:
[0047]
[0048] Among them, , first input the temporal feature, spatial feature, and cross feature into the multi-layer perceptron, and then learn the weight of each feature through the multi-layer perceptron to output the gating vector.
[0049] Finally, achieve multimodal dynamic fusion by introducing a triple-weight gating mechanism. This mechanism further optimizes the feature integration on the basis of spatio-temporal fusion, and the specific formula is as follows:
[0050]
[0051] Among them, are the gating weights for temporal, spatial, and cross features respectively.
[0052] This integration process ensures the full fusion of temporal features, spatial features, and cross features, providing strong support for fault prediction and location. Through this method, it is possible to more accurately analyze and respond to various inter-service interaction and communication problems.
[0053] S2-6. Fault prediction and location, output the fault location result.
[0054] In the fault prediction and location system, first, global average pooling is performed on the fused spatio-temporal features. This step helps reduce the dimension of the data while retaining key spatio-temporal information, thus providing a more concise and effective feature representation for subsequent analysis and obtaining features .
[0055] Next, the features , through a classifier, realize system-level fault prediction. This classifier is based on the processed average pooling features and can effectively identify various possible fault modes in the system. The working principle of the classifier is based on supervised learning, and the model is trained with known fault and normal operation data to quickly and accurately predict faults during actual operation.
[0056] In addition, the features , by classifying node features, can further realize service-level fault location.
[0057] S3. During the process of training the fault location model, a binary marking method is adopted to distinguish service states.
[0058] S4. Automated testing is introduced into the real-time running microservice system. In addition, the trained fault location model is integrated into the online microservice system to predict and locate potential faults in real time.
[0059] Advantages of the present invention:
[0060] (1) Enhance the spatio-temporal feature expression ability: The present invention proposes a Transformer-BiLSTM model that fuses spatio-temporal features. This model adopts time-aware attention and dynamic graph attention combined with spatio-temporal joint decay factors, and through multi-modal adaptive fusion technology, effectively captures the dependence and competition relationships between service performance resources. This innovative design enables the model to adapt to the complexity and dynamics of the microservice architecture, significantly improving the microservice anomaly detection and fault location effects based on the dependence relationships and link data of microservices.
[0061] (2) Dynamically adjust the number of service instances: In the online prediction system, the present invention can dynamically adjust the number of service instances of each service in the microservice system according to the results predicted by the model. This adaptive adjustment mechanism not only ensures the continuous and stable operation of the microservice system in the face of potential fault risks, but also significantly improves the reliability and robustness of the system. Description of the Drawings
[0062] Figure 1 Flow chart of Transformer-BiLSTM for fusing spatio-temporal features
[0063] Figure 2 It is an overview of potential fault prediction and fault location for microservices. Specific implementation manners
[0064] The present invention will be further described below in conjunction with the accompanying drawings.
[0065] In view of the above problems, the present invention proposes a microservice application fault location method based on link data learning, as Figure 1 shown. This method uses a Transformer-BiLSTM model that fuses spatio-temporal features to predict and locate faults in the microservice system. To meet the requirements of the actual application scenario, this model is integrated into the microservice system to achieve real-time fault prediction and location functions for online applications.
[0066] Figure 2 An overview of the microservice fault detection and location process is provided, which is divided into an offline training stage and an online prediction stage. In the offline training stage, the microservice application is deployed in a test environment, and the prediction model is trained by collecting application-level log data. To fully train the model, not only the log data when the service is running normally is collected, but also error log information is included. By combining the methods of fault injection and automated testing, the execution logs when the application goes wrong are simulated, and these tests are designed to simulate the calls of microservices and user requests. To increase the diversity of experimental data, different service configurations are also set for testing. The online prediction stage focuses on monitoring the running state of the application in the production environment to predict potential errors and their specific locations, including which microservice the fault occurs in and the type of the fault. In this stage, the system continuously collects and analyzes the log data from the running application, extracts features in a similar way to data annotation, and uses the trained model to predict potential faults.
[0067] S1. To simulate the real usage scenario, the present invention performs mixed fault injection on two open-source microservice systems, TrainTicket and Shop-Sock, and then uses the Istio service mesh of Kubernetes to collect microservice link data through its distributed tracing and preprocess the data.
[0068] Further: The step 1 includes the following sub-steps:
[0069] S1-1. Perform mixed fault injection on two open-source microservice systems, TrainTicket and Shop Sock.
[0070] To simulate the real application scenario and test the robustness, fault tolerance and recovery ability of the microservice architecture, the present invention adopts the following four methods.
[0071] Utilize the traffic control function of Istio service mesh in Kubernetes to simulate network latency and packet loss phenomena.
[0072] Set limits on the memory or CPU usage of containers to simulate resource exhaustion scenarios.
[0073] Implement exception handling through code injection technology to simulate service crashes.
[0074] Use the fault injection function of Istio to simulate HTTP error codes, such as 500, 502, etc.
[0075] Through the combination of the above strategies, the present invention can comprehensively simulate network problems, resource exhaustion, service crashes, and different types of system failures, thereby providing test data for further anomaly detection and analysis.
[0076] S1-2: Use the Istio service mesh of Kubernetes to collect and process microservice link data through its distributed tracing.
[0077] After generating historical error data using the fault injection technology, the present invention obtains the link data of microservices through the service mesh (Istio), including metrics such as CPU usage rate, memory consumption, and network traffic, and constructs a dataset based on these data. When running the microservice system in the simulation scenario, abnormal behaviors often occur, such as abnormal increases in CPU and memory usage rates, increased network latency, and service call timeouts. By comparing normal and abnormal data, reasonable thresholds are set for each performance metric in this paper, and a training set for model training is sorted and generated.
[0078] S1-3: Data preprocessing.
[0079] First, define the link data matrix as , where T is the time step, N is the number of microservices, and M is the feature dimension (including CPU usage rate, memory consumption, network traffic, request return value, request response time, and one-hot encoded service name).
[0080] Next, construct a call relationship graph between services and set a static service call adjacency matrix . If service i calls service j, then , otherwise it is 0.
[0081] To meet the requirements of model training, all numerical data is standardized using MinMaxScaler, and the service name is transformed into a vector form through one-hot encoding , where K is the total number of services.
[0082] S2: Construction of a fault prediction and location model.
[0083] After data preprocessing, the present invention performs fault prediction and location of microservice applications through a Transformer-BiLSTM model that fuses spatio-temporal features. This model combines spatio-temporal features and can fully learn the temporal dependence relationships and spatial structure features in the microservice system.
[0084] S2-1, Temporal embedding and causal position encoding.
[0085] In the microservice system, the propagation of faults has obvious causal temporal characteristics. For example, when service M fails, the associated service N may also exhibit anomalies. To accurately capture and understand this causal relationship, the causal position can be encoded, and the specific formula is as follows:
[0086] ;
[0087] where, learns the causal weight through historical anomaly data, represents 1 if the time is the potential cause of time , otherwise 0, is the standard sine position encoding vector.
[0088] In addition, the dynamic causal injection technology is also used to simulate and observe the response of the microservice system in the face of configuration changes, traffic fluctuations, changes in service calls, etc. Through this technology, the dynamic changes of the system can be observed in real time, and the causal relationship model of the present invention can be further verified and adjusted. The specific formula is as follows:
[0089] ;
[0090] where, is the input feature matrix, is the embedding layer weight matrix.
[0091] S2-2, Time-aware multi-head attention mechanism.
[0092] In the present invention, to enhance the model's ability to extract temporal features from the input sequence, a time-aware multi-head attention mechanism is introduced. This mechanism specifically considers the spatial relationship of the service call path and enhances the model's spatio-temporal perception ability by integrating the service call path distance into the attention calculation.
[0093] First, by combining the service space topology relationship graph and the time decay factor, a spatio-temporal joint decay factor is constructed. This decay factor not only considers the time factor but also incorporates the spatial distance between services, thus achieving the joint perception of spatio-temporal information. The specific formula is as follows:
[0094] ;
[0095] Among them, is the shortest path distance from service i to j, 、 and are learnable parameters; and represent the timestamps of service i and service j.
[0096] Next, the above spatio-temporal joint decay factor is used to optimize the attention mechanism, and the specific formula is as follows:
[0097] ;
[0098] Among them, Q, K, and V are the query, key, and value respectively.
[0099] Finally, this optimized attention mechanism is extended to the multi-head setting to enhance the model's expressive power and generalization ability. Each head processes different representation subspaces, thereby capturing multi-faceted features of the input data. The final output of the multi-head attention is the concatenation of the outputs of each head, and the specific formula is as follows:
[0100] ;
[0101] ;
[0102] Among them, is the projection matrix after concatenating the multi-head outputs, 、 and are the projection matrices of the k-th attention head.
[0103] S2-3. Temporal modeling of BiLSTM.
[0104] The temporal features of the output of the Transformer are input into the BiLSTM network to capture local temporal dependencies. BiLSTM performs temporal modeling through forward and backward paths, and finally obtains the final temporal features by concatenating the bidirectional hidden states. The specific formula is as follows:
[0105] ;
[0106] Among them, and are the hidden states of the forward and backward LSTMs, is the output temporal feature.
[0107] S2-4. Spatial feature extraction.
[0108] When performing spatial feature extraction, first construct node features and map the temporal features into the graph space. This process involves calculating the attention scores of each node to its neighbors in order to more accurately capture the interactions and information flow between nodes. The specific formula is as follows:
[0109] ;
[0110] Next, to solve the multi-service collaborative detection problem, a spatio-temporal hypergraph attention mechanism is adopted. This mechanism can reflect the collaborative anomalies of multiple services. Each service and all its upstream dependent services form a hyperedge , is the set of direct upstream services of service i, is the hyperedge of service i. The specific formula is as follows:
[0111] ;
[0112] Among them, is the max pooling operation, is the activation function, is the attention coefficient vector.
[0113] Finally, to solve the complex relationships between services in the microservice architecture, the dependency relationships are expressed through hyperedge feature aggregation, thereby helping to detect service anomaly problems. The specific formula is as follows:
[0114] ;
[0115] S2-5, Multimodal Adaptive Fusion.
[0116] When dealing with a large amount of data generated by inter-service communication, such as log files, performance metrics, and network traffic, etc., a multimodal adaptive gating mechanism is proposed to effectively analyze this data. A triple gating system is designed to optimize the fusion weights for temporal features, spatial features, and cross features.
[0117] First, to improve the accuracy of the model, project the cross-attention features and calculate a comprehensive cross-attention mechanism. This mechanism involves the application of the attention weight matrix. The specific formula is as follows:
[0118] ;
[0119] Among them, , and are the attention weight matrices, is the temporal feature, is the spatial feature, is a cross feature, is the feature dimension.
[0120] Next, in order to enable the model to effectively extract and utilize information from each modality, a multi-modal gating vector is introduced. This vector helps the model adjust the weights of information according to the characteristics of different data, so as to perform anomaly detection more effectively. The specific formula is as follows:
[0121] ;
[0122] Among them, (constraining the sum of gating to be 1), first input the time feature, space feature and cross feature into the multi-layer perceptron, and then learn the weights of each feature through the multi-layer perceptron to output the gating vector.
[0123] Finally, a three-weight gating mechanism is introduced to achieve multi-modal dynamic fusion. This mechanism further optimizes feature integration on the basis of spatio-temporal fusion. The specific formula is as follows:
[0124] ;
[0125] Among them, are the gating weights of time, space and cross features respectively.
[0126] This integration process ensures the full fusion of time features, space features and cross features, providing strong support for fault prediction and location. Through this method, various service interactions and communication problems can be analyzed and responded to more accurately.
[0127] S2-6, Fault Prediction and Location.
[0128] In the fault prediction and location system of the present invention, first, global average pooling is performed on the fused spatio-temporal features. This step helps reduce the dimension of the data while retaining key spatio-temporal information, thus providing a more concise and effective feature representation for subsequent analysis. The specific formula is as follows:
[0129] ;
[0130] Among them, N is the total number of services, i represents the current microservice, is the feature vector after global pooling.
[0131] Next, a system-level fault prediction is achieved through a carefully designed classifier. This classifier is based on the processed average pooling features and can effectively identify various possible fault modes in the system. The working principle of the classifier is based on supervised learning. The model is trained with known fault and normal operation data, so as to quickly and accurately predict faults during actual operation. The specific formula is as follows:
[0132] ;
[0133] Among them, are system-level classifier parameters.
[0134] In addition, by classifying node features, fault location at the service level can be further achieved. The specific formula is as follows:
[0135] ;
[0136] Among them, are node-level classifier parameters.
[0137] S3. During the process of training the fault prediction model, a binary marking method is adopted to distinguish service states.
[0138] Specifically, if the performance metrics or link data of a service show anomalies, it indicates that the service has a fault. At this time, the status is marked as 1; conversely, if all performance metrics are within the normal range, the status is marked as 0, indicating that the service is running normally. For the training of the fault classification model, a more detailed marking system is used to identify specific fault types. If the CPU usage rate is too high, it is marked as 1, memory leakage is marked as 2, and when network latency occurs, it is marked as 3; if the request return value and request response time are abnormal, it is marked as 4, indicating multi-instance faults. If all performance metrics of the service are within the normal range, the instance status is marked as 0, indicating no fault. This marking strategy not only helps the model accurately learn and predict service states, but also effectively identifies and classifies specific fault types, thus providing guidance for subsequent fault handling and optimization.
[0139] S4. To enhance the robustness of the system and simulate real fault scenarios, automated testing is introduced into the real-time running microservice system. This enables the present invention to simulate various potential fault situations in a controlled and safe environment. In addition, the pre-trained model is integrated into the online microservice system to predict and locate potential faults in real time. The main function of this model is to use the current system state and historical data to predict the probability of each microservice having a fault.
[0140] When the model predicts that a certain microservice may have a fault, the system will automatically increase the number of instances of this service. This adaptive response mechanism ensures that even when facing the risk of potential faults, the microservice application can maintain stable operation. Through this strategy, not only the system's fault prevention ability is improved, but also the dynamic allocation of resources is optimized, thus enhancing the availability and reliability of the service. This integrated automated testing and real-time fault prediction system is the key to ensuring the continuous and stable operation of the microservice architecture.
[0141] Table 1
[0142]
[0143] Table 2
[0144]
[0145] Table 1 and Table 2 respectively represent the performance of fault prediction accuracy and Top-k prediction accuracy. The results show that in the ShopSock system, the present invention exhibits better performance compared to the TrainTicket system. This difference mainly comes from the relatively small number of services in the ShopSock system and the more concise service execution traces. Compared with the CNN-LSTM and LSTM methods, the present invention performs better in predicting potential faults and can more effectively adapt to the dynamic changes and complexity of the microservice system.
[0146] Table 3
[0147]
[0148] Table 4
[0149]
[0150] Table 3 shows the comparison results of the present invention with other methods in terms of fault type prediction accuracy. It can be seen from the table that in both the ShopSock and TrainTicket systems, the method proposed by the present invention shows significant performance advantages compared to CNN-LSTM and LSTM. This advantage mainly stems from the in-depth analysis of the spatio-temporal characteristics of link data and service dependencies by the present invention, thus greatly improving the accuracy of fault type identification.
[0151] Table 4 shows the fault prediction accuracy in the actual application scenario. The results show that the TrainTicket system performs relatively well in predicting configuration faults. However, due to its extensive microservice architecture resulting in complex dependencies and interactions, the accuracy of predicting potential multi-instance faults is relatively low. In contrast, the Sock Shop system shows higher prediction accuracy and recall rate in configuration faults (such as memory leaks and CPU usage). In addition, the Sock Shop system is also significantly superior to the TrainTicket system in the prediction accuracy and recall rate of multi-instance faults. This difference may be due to the smaller number of microservices in the Sock Shop system and the clearer service dependencies, which helps the training and prediction of the fault detection model.
[0152] The present invention proposes an innovative method that focuses on learning from link data to achieve error prediction and fault location for microservice applications. Through a series of experimental verifications, the present invention can achieve high accuracy in the prediction within the application of potential errors, faulty microservices, and fault types, and is superior to the state-of-the-art fault diagnosis methods for distributed systems. In addition, the present invention can also effectively predict potential errors caused by actual fault cases, providing strong support for improving the stability and reliability of microservice systems.
Claims
1. A method for locating faults in microservice applications by learning from link data, characterized in that: The following steps are involved: S1. Perform hybrid fault injection on the open source microservice system, and then use Kubernetes' Istio service grid to collect service link data through its distributed tracing and pre-process the data. S2. Build a fault location model: After data preprocessing, the Transformer-BiLSTM model that integrates spatiotemporal features is used to locate the faults of microservice applications. The specific implementation process is as follows: S2-1. Perform temporal embedding and causal position encoding on the preprocessed data. The specific process is as follows: Encode the causal position, the specific formula is as follows: P causal (t)=P(t)+∑ τ<t a t,τ ·CausalLink(τ→t); Among them, α t,τ Causal weights are learned through historical abnormal data. CausalLink(τ→t) means that if time τ is the potential cause of time t, it is 1, otherwise it is 0. P(t) is the standard sinusoidal position encoding vector. Dynamic causal injection technology is used to simulate and observe the response of the microservice system in the face of configuration changes, traffic changes, and changes in service calls. The specific formula is as follows: Among them, X is the input feature matrix, W e is the embedding layer weight matrix; S2-2. After encoding, a time-aware multi-head attention mechanism is introduced to consider the spatial relationship of the service call path. By integrating the service call path distance into the attention calculation, the spatiotemporal perception ability is enhanced. S2-3. Based on the output of the time-aware multi-head attention mechanism, BiLSTM time series modeling is performed to obtain the final time features; S2-4, combine the spatiotemporal hypergraph attention mechanism to extract spatial features; S2-5, based on spatial features, multi-modal adaptive fusion is performed through a multi-modal adaptive gating mechanism to obtain features S2-6, perform fault prediction and location based on the multi-modal adaptive fusion results, and output the fault location results: perform global average pooling on the fused spatiotemporal features to obtain the feature h global ; A classifier is used to achieve system-level fault prediction. In addition, based on the feature Service-level fault location is achieved by classifying node features; S3. In the process of training the fault location model, a binary labeling method is used to distinguish the service status and output the fault location result; S4. Automated testing is introduced into the real-time microservice system. In addition, the trained fault location model is integrated into the online microservice system to predict and locate potential faults in real time. When the model predicts that a service will fail, the number of instances of the service is automatically increased.
2. The method for locating a fault in a microservice application by learning from link data according to claim 1, characterized in that: The specific implementation process of step S1 is as follows: S1-1. Hybrid fault injection into open source microservice systems: Use the traffic control of the Istio service grid in Kubernetes to simulate network delay and packet loss; set limits on the memory or CPU usage of containers to simulate resource exhaustion scenarios; implement exception handling through code injection technology to simulate service crashes; use Istio's fault injection function to simulate HTTP error codes; S1-2. Use Kubernetes' Istio service grid to collect service link data and process the data through its distributed tracing: After the fault injection technology generates historical error data, the service link data, including CPU usage, memory consumption, and network traffic, is obtained through the service mesh Istio, and a data set is built based on this data. By comparing normal and abnormal data, thresholds are set for each performance indicator, and a training set is generated. S1-3. Data preprocessing: First, the link data matrix is defined as Where T is the time step, N is the number of services, and M is the feature dimension, including CPU usage, memory consumption, network traffic, request return value, request response time, and service name after one-hot encoding; Then, build the call relationship graph between services and set the static service call adjacency matrix If service i calls service j, then A ij =1, otherwise 0; all numerical data are standardized using MinMaxScaler, and service names are converted into vector form through one-hot encoding K is the total number of services.
3. The method for locating a fault in a microservice application by learning from link data according to claim 2, characterized in that: The specific implementation process of the time-aware multi-head attention mechanism is as follows: First, by combining the service space topology graph and the time decay factor, a spatiotemporal joint decay factor is constructed. The specific formula is as follows: Where SPD(i,j) is the shortest path distance from service i to j, α, β and γ are learnable parameters; t i and t j Represents the timestamps of service i and service j; The above-mentioned spatiotemporal joint attenuation factor is used to optimize the attention mechanism. The specific formula is as follows: Among them, Q, K and V are query, key and value respectively; Finally, the optimized attention mechanism is extended to a multi-head setting, where each head processes a different representation subspace to capture multi-faceted features of the input data; the final output of the multi-head attention is the concatenation of the outputs of each head.
4. The method for locating a fault in a microservice application by learning from link data according to claim 3, characterized in that: The specific implementation process of BiLSTM time series modeling in step S2-3 is as follows: The link data of the service is processed through Transformer, and then the high-dimensional features that can reflect the dynamics of the time series are input into the BiLSTM layer. The time series features output by Transformer It is input into the BiLSTM network to capture local time dependencies; BiLSTM performs time series modeling through forward and reverse paths, and finally obtains the final time feature H by concatenating the bidirectional hidden states. time .
5. The method for locating a fault in a microservice application by learning from link data according to claim 4, characterized in that: The specific implementation process of the spatiotemporal hypergraph attention mechanism is as follows: When performing spatial feature extraction, the spatiotemporal hypergraph attention mechanism is first used. Each service and all its upstream dependent services form a hyperedge ε i ={i}∪{j|∈Upstream(i)}, where Upstream(i) is the set of direct upstream services of service i, ε i is the hyperedge of service i, and the specific formula is as follows: MaxPool(·) is the maximum pooling, LeakyReLU(·) is the activation function, and a is the attention coefficient vector; Then, the dependency is expressed through hyperedge feature aggregation. The specific formula is as follows: H space =∑ ε α i,ε ·MeanPool({h j |j∈ε})。 6. The method for locating a fault in a microservice application by learning from link data according to claim 5, characterized in that: The specific implementation process of multi-modal adaptive fusion through the multi-modal adaptive gating mechanism is as follows: A multi-modal adaptive gating mechanism is proposed, and a triple gating system is designed to optimize the fusion weights of temporal features, spatial features, and cross-features. First, the criss-cross attention features are projected and a comprehensive criss-cross attention mechanism is calculated: Among them, W Q , W K and W V is the attention weight matrix, H time is the time characteristic, H space is the spatial feature, H cross is the cross feature, d is the feature dimension; Next, the multimodal gating vector is introduced. The specific formula is as follows: g t ,g s ,g c =softmax(MLP([H time ||H space ||H cross ])); Among them, g t +g s +g c = 1, firstly input the time feature, spatial feature and cross feature into the multi-layer perceptron, then learn the weight of each feature through the multi-layer perceptron and output the gate vector; Finally, multi-modal dynamic fusion is achieved by introducing a triple-weight gating mechanism. The specific formula is as follows: H fusion =g t ·H time +g s ·H space +g c ·H cross ; Among them, g t ,g s ,g c They are the gating weights of time, space, and cross features respectively.
7. The method for locating a fault in a microservice application by learning from link data according to claim 6, characterized in that: The binary marking method is specifically used to distinguish the service status: if the performance indicator or link data of the service shows abnormality, indicating that the service fails, the status is marked as 1; On the contrary, if all performance indicators are within the normal range, the status is marked as 0, indicating that the service is running normally; For the training of the fault classification model, a more detailed marking system is used to identify the specific fault type. If the CPU usage is too high, it is marked as 1, if the memory leak is marked as 2, and if a network delay occurs, it is marked as 3; if the request return value and request response time are abnormal, it is marked as 4, indicating multiple instance failures; if all performance indicators of the service are within the normal range, the instance status is marked as 0, indicating no faults.
Citation Information
Patent Citations
Microservice call chain anomaly detection method and device based on graph convolutional neural network
CN115185736A