An RPA Fault Diagnosis Method, Device, and Storage Device Based on Hypergraph
By constructing heterogeneous instance relationship hypergraphs and using heterogeneous hypergraph attention neural networks, the problem of inaccurate instance-level fault diagnosis in microservice systems is solved, and fast and accurate fault location and fault category identification are achieved.
Patent Information
- Application Number
- CN202510317450.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-18
AI Technical Summary
The prior art is difficult to locate the root cause of instance-level failure in microservice systems, and ignores the heterogeneous relationship between service instances, resulting in inaccurate fault diagnosis.
By constructing a heterogeneous instance relationship hypergraph and combining a heterogeneous hypergraph attention neural network, dynamic microservice instance fault diagnosis can be achieved, and the root cause instances and fault categories of failures can be screened out from a large number of microservice instances.
It realizes rapid and accurate detection of microservice system failures, locates the root cause instances and fault categories, and adapts to the dynamic changes of instances in the microservice system.
Smart Images

Figure CN119862063B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of RPA fault diagnosis based on instance heterogeneous relationship hypergraphs, and particularly to a hypergraph-based RPA fault diagnosis method, device, and storage device. Background Art
[0002] As a distributed architecture, the microservices architecture is popular among cloud-native enterprises because of features such as resource flexibility, service loose coupling, and lightweight deployment. In a microservices system, services jointly complete user requests through mutual calls. Services are dynamically deployed in containers in the form of instance replicas and run on shared hosts. Although this unique service running method brings good resource flexibility and service loose coupling to the microservices system, it also makes large-scale microservices systems have complex dependency relationships and dynamic structures. When a microservices system fails, it is precisely this dependency and dynamicity that the failure of a service may spread along with the service call relationship and the shared host. If the root cause of the failure is not diagnosed in time and mitigation measures are not taken, it may cause performance degradation or process termination of services in a large range, seriously affecting the user experience.
[0003] Maintaining the observability of the microservices system is a prerequisite for diagnosing faults. Some microservices systems integrate an RPA system to initially automatically process monitoring data, and the RPA system reminds engineers by sending out anomaly alerts. During the operation of an instance, due to reasons such as metric anomalies, latency anomalies, and business anomalies, when the relevant monitoring data shows anomalies, the RPA system processes them and triggers the generation of an RPA event <TimeStamp, Pod_Name, Anomaly>, which records the anomaly alert of a certain instance at a certain moment. When a microservices system fails, a large number of RPA events are often generated, especially during the process of the failure spreading in the microservices system. However, manually extracting effective fault information and performing fault diagnosis often requires a large amount of time and effort. Traditional heuristic methods require a large amount of diagnostic time, which is unacceptable for a system that urgently needs to restore availability from service degradation or even interruption risks. Therefore, it is crucial to design an accurate and efficient automated fault diagnosis method to locate the root cause service instance and fault category of the fault.
[0004] Currently, a large number of studies are based on observable monitoring data to automatically diagnose the root cause of faults and fault categories, so as to maintain the availability of the microservices system. However, these methods do not consider the dynamicity of service instances, so they cannot locate the root cause of instance-level faults. At the same time, these methods ignore the heterogeneous relationships between service instances and it is difficult to accurately locate the root cause instance. Summary of the Invention
[0005] To solve the problems of inability to locate the root cause of instance-level faults and difficulty in accurately locating the root cause instances, the present invention proposes a hypergraph-based RPA fault diagnosis method, device, and storage device, which are used to automatically, quickly, and accurately locate the root cause service instances and fault categories of faults. The present invention constructs a heterogeneous instance relationship hypergraph, combines a heterogeneous hypergraph attention neural network to fully represent and mine diverse propagation patterns of faults, and realizes dynamic microservice instance fault diagnosis through an instance-agnostic neural network, which can screen out the root cause instances and fault categories of faults from a large number of microservice instances to assist engineers in fault diagnosis.
[0006] The present invention provides a hypergraph-based RPA fault diagnosis method, which mainly includes:
[0007] S1: Construct a heterogeneous instance relationship hypergraph based on trace information, service information, and deployment information;
[0008] S2: Group RPA events according to each instance on the heterogeneous instance relationship hypergraph and encode them into RPA event expressions of the instances;
[0009] S3: On the basis of step S2, update and aggregate the instance node features of the heterogeneous instance relationship hypergraph based on the heterogeneous hypergraph attention neural network to obtain instance-level fault features and graph-level fault features;
[0010] S4: Based on the instance fault features and graph-level fault features, jointly learn the root cause localization and fault classification tasks, and finally obtain the ranking of the root cause instances of the faults and the fault types.
[0011] A storage device stores instructions and data for implementing a hypergraph-based RPA fault diagnosis method.
[0012] A hypergraph-based RPA fault diagnosis device includes: a processor and a storage device; the processor loads and executes the instructions and data in the storage device for implementing a hypergraph-based RPA fault diagnosis method.
[0013] The beneficial effects brought by the technical solution provided by the present invention are: Based on the instance-based RPA events of the present invention, the heterogeneous relationships of the instances are organized through a hypergraph, and the faults of the microservice system are diagnosed through a heterogeneous hypergraph attention neural network. 1) When a fault occurs in the microservice system, it can be quickly and accurately detected, and the root cause instances and fault categories of the faults are given; 2) The heterogeneous relationships between instances in the microservice system are comprehensively organized through the hypergraph; 3) The propagation of faults in the microservice system is captured through the heterogeneous hyper attention neural network, and an instance-agnostic neural network model is realized in combination with a root cause scorer and a fault classifier, which can well adapt to the dynamic changes of instances in the microservice system. Description of the Drawings
[0014] The present invention will be further described below in conjunction with the accompanying drawings and embodiments. In the accompanying drawings:
[0015] Figure 1 is a flowchart of a hypergraph-based RPA fault diagnosis method in an embodiment of the present invention;
[0016] Figure 2 is a schematic diagram of constructing a heterogeneous instance relationship hypergraph based on tracking information, service information, and deployment information in an embodiment of the present invention;
[0017] Figure 3 is the loss situation during training by the stochastic gradient descent algorithm in an embodiment of the present invention;
[0018] Figure 4 is a schematic diagram of the operation of hardware devices in an embodiment of the present invention. Specific Embodiments
[0019] For a clearer understanding of the technical features, objectives, and effects of the present invention, the specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0020] As Figure 1 shown, the present invention proposes a hypergraph-based RPA fault diagnosis method, including:
[0021] S1: Construct a heterogeneous instance relationship hypergraph based on tracking information, service information, and deployment information. Specifically:
[0022] S1.1: Collect service information, deployment information, and call information in the microservice system: Service information means that in the microservice system, a microservice has one or more microservice instances. For example, Figure 2 in (a) of Figure 2 , Service A has three instances a1, a2, and a3; Deployment information means that in the microservice system, one or more microservice instances are deployed and run on the same host. For example, Figure 2 in (b) of
[0023]
[0024] Build a service hyperedge: All instances of the same service share the code of this service and place them in the same hyperedge. Figure 2 As shown in (d), the three instances of service A, a1, a2 and a3, are within the same service hyperedge.
[0025] Build a deployment hyperedge: All instances on the same host share the resources of this host and place these instances on the same hyperedge. Figure 2 As shown in (e), the three instances a1, b1 and c2 on host h1 are in the same deployment hyperedge.
[0026] Construct a calling hyperedge: All calling instances that call the same called instance share the state of the called instance, and place these called instances and the calling instance on the same hyperedge. Figure 2 As shown in (f), all call instances b1, b2, c1 and c2 of call a1 and instance a1 are in the same call hyperedge.
[0027] Construct an instance heterogeneous relationship hypergraph: put all service hyperedges, deployment hyperedges, and call hyperedges in the same hypergraph to obtain an instance heterogeneous relationship hypergraph .like Figure 2 As shown in (g), the instance heterogeneous relationship hypergraph consists of 7 instance nodes and 3 hyperedges.
[0028] S2: Group the RPA events according to each instance on the heterogeneous instance relationship hypergraph and encode them into RPA event expressions of the instances. Specifically:
[0029] S2.1: For all instance nodes on the heterogeneous instance relationship hypergraph, collect the RPA events of each instance.
[0030] During the instance operation, due to abnormal indicators, abnormal delays, abnormal business, etc., the relevant monitoring data will automatically generate RPA events.<TimeStamp,Pod_Name,Anomaly> , records the abnormal alarm of a certain instance at a certain time. For example, the RPA event <1684475>398705,a1,CPU_usage_high> records that the CPU usage of instance a1 was too high at the UNIX timestamp 1684475398705.
[0031] S2.2: Group and sort all RPA events according to instances, and encode the RPA event sequence into the RPA event expression of the instance.
[0032] All RPA events are grouped according to the instance and sorted according to the time of occurrence to obtain the RPA event sequence of the instance:
[0033]
[0034] Among them, represents the RPA event sequence of the instance nodes in the instance heterogeneous relationship hypergraph ; represents the -th RPA event of the instance node ; represents the -th RPA event of the instance node ; represents the -th RPA event of the instance node
[0035] Regarding the RPA event sequence as a sentence in natural language, and the RPA event as a word in the sentence, vector encoding is performed on through the FastText word embedding model to obtain the high-dimensional representation of the encoded RPA event sequence .
[0036] S3: Update and aggregate the features of the nodes in the heterogeneous instance relationship hypergraph based on the heterogeneous hypergraph attention neural network to obtain the instance-level fault representation and the graph-level fault representation. Specifically:
[0037] S3.1: For the instance heterogeneous relationship hypergraph , denote as the instance node, as the hyperedge, as the set of instance nodes ; as the set of hyperedges . The high-dimensional representation of each instance node is updated by fusing node features through the heterogeneous hypergraph attention neural network. The specific steps are as follows:
[0038] Aggregate the node features to the hyperedges through weighted average to obtain the hyperedge feature :
[0039] ,
[0040] ,
[0041] Among them, represents the features of the instance nodes in the hypergraph, k represents the k-th instance node, represents the j-th node feature, j = 1, 2,... , represents a hyperedge the number of internal nodes, represents the hyperedge feature, represents the encoded high-dimensional representation, that is, the RPA event representation obtained by embedding the RPA event sequence of instance node k;
[0042] Connected by an equation, it means that the initial RPA event representation is assigned to the instance node as the input feature, that is, as the initial node failure representation, which is also the instance node feature in the hypergraph ; In the heterogeneous hypergraph attention neural network, it is updated by the adjacent hyperedge feature to , which also means that the initial RPA event representation is also updated, and this update can be considered as the integration of the initial features of the instance node into the features of the adjacent hyperedges.
[0043] Instance-level fault feature , which is the updated RPA event representation . An instance is a node in the heterogeneous instance relationship hypergraph. Therefore, the instance-level fault representation can also be said to be node-level. Fusing all the node-level fault features together is the graph-level fault feature, that is, the graph-level fault feature.
[0044] Select the attention weight according to the hyperedge category and calculate the attention score of instance node and hyperedge : :
[0045]
[0046] Among them, is the type of hyperedge , is the attention weight corresponding to hyperedge , is the weight for linear transformation, represents concatenation, is the activation function.
[0047] Normalize the attention score through the softmax function :
[0048]
[0049] Among them, represents the normalized attention score; is the set of all hyperedges of instance node k, , representing a certain adjacent hyperedge of the updated instance node k represents the instance node k and the updated adjacent hyperedge The normalized attention score.
[0050] Update the instance-level fault feature by weighted averaging the hyperedge features according to the attention score :
[0051]
[0052] Obtain the instance node according to the instance-level fault feature The updated RPA event representation :
[0053]
[0054] Pool the RPA event representations of all nodes to obtain the feature of the hypergraph :
[0055]
[0056] Among them, represents the number of instance nodes in the hypergraph.
[0057] S4: Based on the instance fault representation and the graph-level fault representation, jointly learn the root cause localization and fault classification tasks, and finally obtain the root cause instance ranking and fault type of the fault. Specifically:
[0058] For fault classification, since the fault category remains unchanged, the corresponding fault classifier can be implemented through MLP, denoted as , the formula is as follows:
[0059] ,
[0060] ,
[0061] Receive the hypergraph feature , and output the probability score of a fixed dimension . m is the total number of fault categories, represents the probability of the -th category, = 1, 2,..., m, and select the category with the largest probability score
[0062] For the root cause localization task, considering the characteristics of the dynamic change of the instance node, design a dynamic instance root cause scorer Scorer, denoted as , the one-dimensional root cause score is output, which can score all instance nodes. The formula is as follows:
[0063] ,
[0064] ,
[0065] Among them, represents the updated RPA event expression of the instance node , represents concatenation, represents the graph-level fault feature, k = 1, 2,..., , represents the number of instance nodes in the hypergraph, represents the maximum root cause score;
[0066] Consider making the instance node have a global view. Therefore, use the above formula to concatenate the updated RPA event expression of the instance node with the graph-level fault feature as the input of the Scorer to obtain the root cause score of the instance node . Select the instance with the largest root cause score as the root cause instance.
[0067] Jointly learn the root cause localization task and the fault classification task of the fault. Both the root cause localization and the fault classification tasks are multi-classification tasks. The loss function uses the cross-entropy loss function. Add the loss functions of the two tasks for joint learning. Optimize the entire network model through the stochastic gradient descent algorithm. This network model includes a heterogeneous hypergraph attention neural network, a fault classifier, and a dynamic instance root cause scorer. The stochastic gradient descent algorithm is:
[0068]
[0069] Among them represents the number of fault samples, is the number of instance nodes in the fault sample . is an indicator function that takes 1 when the root cause node in the fault sample i is k. represents the predicted score of node k in sample i. is the fault sample The number of fault categories in. is an indicator function that takes 1 when the j-th category in the fault sample i is the fault category. represents the predicted score of category j in the fault sample i.
[0070] The present invention uses the open-source dataset AIOps-2022 as an example.
[0071] The AIOps-2022 dataset comes from the training set of the AIOps 2022 Challenge. It is obtained by injecting three levels of fault categories, namely hosts, services, and instances, into an e-commerce platform. The AIOps-2022 dataset contains RPA events and heterogeneous instance relationships at the service instance level during the operation of the microservice system, records the duration of the fault time period, as well as the root cause service instance and fault category for each fault time period. The detailed information of the dataset is as follows:
[0072]
[0073] Train a heterogeneous hypergraph neural network, a root cause scorer, and a fault classifier on a training set with 300 fault samples. The heterogeneous hypergraph neural network has 8 attention heads and a 2-layer structure. Each attention head contains 3 attention weights (the same as the types of heterogeneous hypergraphs). The input weight is 128-dimensional, the hidden weight is 64-dimensional, and the output weight is 96-dimensional. The root cause scorer is a fully connected network, consisting of an input layer, three residual blocks, and an output layer. The input dimension is 192, the hidden layer dimension is 15, and the output layer dimension is 1. Each residual block is composed of two fully connected layers. The fault classifier is a fully connected network, including an input layer, a hidden layer, and an output layer. The input dimension is 96, the hidden layer dimension is 29, and the output layer dimension is 9. Train using the stochastic gradient descent algorithm, with a learning rate of 0.005, 1500 training epochs, a training batch size of 300. The situation of the training loss (loss) is as Figure 3 shown. It can be seen that the loss converges stably. Save the model with the minimum loss and use it for evaluation on the test set.
[0074] Evaluate the model on a test set with 70 fault samples. For the root cause localization task, use the top-j accuracy ( ), and the top-3 average accuracy ( ) as evaluation metrics. Quantifies the probability that the top-j instances output by each method actually contain the root cause service instance. For example, for the test sample , there are service instances, types of fault categories, the fault instance is , the fault category is , represents the top-j root cause instance set. When , the formula is as follows:
[0075]
[0076] The evaluation method locates the overall ability of the root cause examples. In practice, the operator often checks the top 3 results. The calculation formula is as follows:
[0077]
[0078] For the fault classification task, weighted average precision, recall, and are used to test the performance, including true positive (TP), false positive (FP), and false negative (FN). The calculation formula is where , . The performance on the test set is as follows:
[0079]
[0080] Please refer to Figure 4 , Figure 4 which is a schematic diagram of the operation of the hardware device according to an embodiment of the present invention. The hardware device specifically includes: an RPA fault diagnosis device 401 based on a hypergraph, a processor 402, and a storage device 403.
[0081] An RPA fault diagnosis device 401 based on a hypergraph: The RPA fault diagnosis device 401 based on a hypergraph implements the RPA fault diagnosis method based on a hypergraph.
[0082] Processor 402: The processor 402 loads and executes the instructions and data in the storage device 403 to implement the RPA fault diagnosis method based on a hypergraph.
[0083] Storage device 403: The storage device 403 stores instructions and data; the storage device 403 is used to implement the RPA fault diagnosis method based on a hypergraph.
[0084] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A hypergraph-based RPA fault diagnosis method, characterized by: include: S1: Build a heterogeneous instance relationship hypergraph based on call information, service information, and deployment information; Specifically: S1.1: Collect service information, deployment information, and call information in the microservice system; S1.2: Construct service hyperedge, deployment hyperedge and call hyperedge according to service information, deployment information and call information; Build a service hyperedge: All instances of the same service share the code of this service and put these instances in the same hyperedge; Build a deployment hyperedge: All instances on the same host share the resources of this host and are placed in the same hyperedge; Construct a calling hyperedge: All calling instances that call the same called instance share the state of the called instance, and place these called instances and the calling instance on the same hyperedge; S1.3: Put all service hyperedges, deployment hyperedges, and call hyperedges into the same hypergraph to obtain an instance heterogeneous relationship hypergraph ; S2: Group the RPA events according to each instance on the heterogeneous instance relationship hypergraph and encode them into RPA event expressions of the instances; specifically: S2.1: For all instance nodes on the heterogeneous instance relationship hypergraph, collect the RPA events of each instance; S2.2: Group and sort all RPA events according to instances, and encode the RPA event sequence into the RPA event expression of the instance; in, Representing instance nodes in instance heterogeneous relationship hypergraphs The RPA event sequence, Represents an instance node No. RPA events, Represents an instance node No. RPA events, Represents an instance node No. RPA events, Representing a sequence of events The number of RPA events in Will As a sentence in natural language, As words in a sentence, FastText word embedding model is used to embed Perform vector encoding and get Encoded event representation ; S3: Based on step S2, the instance node features of the heterogeneous instance relationship hypergraph are updated and aggregated based on the heterogeneous hypergraph attention neural network to obtain instance-level fault features and graph-level fault features; Instance-level failure characteristics The calculation formula is: According to the instance-level fault characteristics Updated RPA event expressions : All The updated RPA event expressions are fused together to obtain graph-level fault features. : in, represents the normalized attention score, is the weight that changes linearly, represents the hyperedge feature, represents a hyperedge, For instance nodes The set of all hyperedges; Represents the number of instance nodes in the hypergraph; S4: Based on instance-level fault features and graph-level fault features, we jointly learn the root cause location and fault classification tasks, and finally obtain the root cause instance ranking and fault type of the fault; The fault root cause localization task and fault classification task are jointly learned using the cross loss function, and the entire network model is optimized using the stochastic gradient descent algorithm. The network model includes a heterogeneous hypergraph attention neural network, a fault classifier, and a dynamic instance root cause scorer.
2. The hypergraph-based RPA fault diagnosis method according to claim 1, characterized in that: In step S3, for the instance heterogeneous relationship hypergraph , For instance nodes A collection of For super edge A collection of each instance node Event expression , node feature update is fused through heterogeneous hypergraph attention neural network. The specific steps are as follows: By weighted averaging, the instance node features in the hypergraph are Converge to the hyperedge and get the hyperedge feature : , , in, represents the instance node feature in the hypergraph, k represents the kth instance node, represents the input feature of the jth node, j=1,2,..., , represents the number of nodes in the hyperedge e, represents the hyperedge feature, express The encoded event expression; Select attention weights according to hyperedge categories and calculate instance nodes With super edge Attention score : in, For super edge The corresponding attention weights are, is the weight that changes linearly, Indicates splicing, is the activation function; Normalize the attention score through the softmax function : in, represents the normalized attention score, is the set of all hyperedges of instance node k, represents an adjacent hyperedge of the updated instance node k, Represents instance node k and its updated adjacent hyperedges Normalized attention scores.
3. The hypergraph-based RPA fault diagnosis method according to claim 2, characterized in that: Step S4 is specifically as follows: For fault classification, since the fault category remains unchanged, the corresponding fault classifier is implemented through MLP, which is recorded as , the formula is as follows: , , Receive graph-level fault characteristics , output probability scores of fixed dimensions , m is the total number of fault categories, Indicates The probability of the categories, =1,2,...,m, maximum probability score Corresponding categories is the fault category; For the root cause location task, considering the dynamic changes of instance nodes, a dynamic instance root cause scorer Scorer is designed, denoted as , outputs a one-dimensional root cause score, which is used to score all instance nodes. The formula is as follows: , , in, Represents an instance node Updated RPA event expression, Indicates splicing, represents the graph-level fault features, k=1,2,..., , represents the number of instance nodes in the hypergraph, represents the maximum root cause fraction; Consider making the instance node With a global view, the above formula is used to convert the instance node Updated RPA event expressions Fault characteristics at the graph level Splicing, as the input of Scorer, to get the instance node Root cause score , select the instance with the largest root cause score As a root cause example.
4. The hypergraph-based RPA fault diagnosis method according to claim 1, characterized in that: The stochastic gradient descent algorithm is: in, represents the number of fault samples, Indicates a fault sample The number of instance nodes in , Represents the indicator function. When the root cause node in fault sample i is k Take 1, when the j category in the fault sample i is the fault category Take 1, represents the prediction score of node k in fault sample i, For fault samples The number of fault categories in represents the prediction score of category j in fault sample i, Represents the loss value calculated for all samples.
5. A storage device, characterized in that: The storage device stores instructions and data for implementing the hypergraph-based RPA fault diagnosis method described in any one of claims 1 to 4.
6. A hypergraph-based RPA fault diagnosis device, characterized in that: include: A processor and a storage device; the processor loads and executes instructions and data in the storage device to implement the hypergraph-based RPA fault diagnosis method described in any one of claims 1 to 4.
Citation Information
Patent Citations
method for constructing an adaptive decomposable partial repetition code based on a hypergraph
CN109522150A
Microservice system root cause positioning method and device based on graph neural network model
CN117560275A