Fault determination method and device of micro-service system, product and electronic equipment

By constructing an anomaly event cause-effect graph for a microservice system and utilizing a pre-trained model and a multi-agent collaborative architecture, the problem of low accuracy in fault determination in microservice systems is solved, achieving efficient and accurate fault root cause localization and diagnosis.

CN120909878AActive Publication Date: 2025-11-07JINAN INSPUR DATA TECH CO LTD

Patent Information

Application Number
CN202511431214.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-09
Publication Date
2025-11-07
Estimated Expiration
2045-10-09

AI Technical Summary

Technical Problem

Existing shallow feature analysis methods based on microservice systems have low accuracy in fault identification and make it difficult to accurately trace the root cause.

Method used

By acquiring abnormal monitoring information from the microservice system, an abnormal event flow is constructed and an event cause-effect graph is determined. Using a pre-trained model-driven agent, combined with the event cause-effect graph and pre-defined fault modes, abnormal services, propagation paths, and fault classifications are determined. A multi-agent collaborative architecture is then used for fault analysis.

Benefits of technology

It improves the accuracy and efficiency of microservice fault diagnosis, realizes closed-loop diagnosis from anomaly discovery to root cause localization, has stronger cross-domain adaptability and interpretability, and adapts to the dynamically changing microservice environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120909878A_ABST
    Figure CN120909878A_ABST
Patent Text Reader

Abstract

The invention discloses a fault determination method and device for a micro-service system, a product and electronic equipment. Relates to the technical field of artificial intelligence. The method comprises the following steps: determining an abnormal event flow according to abnormal monitoring information; an abnormal event causal graph is determined based on an abnormal event flow, and then a fault point, a fault chain and a fault classification are determined by utilizing an intelligent agent driven by a pre-training model and combining the event causal graph and a fault mode standard. According to the method, anomaly detection, fault classification and root cause positioning of the micro-service system are realized, and closed-loop diagnosis from anomaly discovery to root cause positioning is realized; secondly, performing depth feature engineering and semantic abstraction on the refined data from the perspective of events and causal relationships, fusing system behaviors and dependency relationships dispersed in different modal data into a unified event causal graph through a graph modeling technology, and providing comprehensive and high-dimensional input for fault reasoning of an intelligent agent; and the accuracy of micro-service fault determination is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, in particular to a fault determination method and device for a micro-service system, a product and an electronic device. BACKGROUND

[0002] Micro-service architecture is a software architecture style that splits monolithic applications into small service units deployed independently. With the wide application of micro-service architecture in cloud computing systems, the complexity and dynamics of the system have increased significantly, leading to frequent system failures.

[0003] In related technologies, the method of Artificial Intelligence for IT Operations (AIOps) is used for fault diagnosis. By extracting shallow statistical features such as the number of occurrences of specific error logs, the average length of service call chains, and the fluctuation range of resource usage, these features are compared with historical normal baselines or input into simple models for anomaly recognition. If the features deviate from the normal range, it is determined that there is a potential fault, and the associated nodes are preliminarily located in combination with service dependency relationships. Since these shallow statistical features only reflect the quantitative rules on the surface of the data, they cannot capture deep information, making it difficult to accurately trace the root cause when facing complex faults.

[0004] Therefore, how to improve the accuracy of micro-service fault determination is a technical problem that needs to be solved by those skilled in the art. SUMMARY

[0005] The purpose of the present application is to provide a fault determination method, device, product and electronic device for a micro-service system, to solve the problem of low accuracy in analyzing the fault condition of a micro-service system based on shallow features of the micro-service system.

[0006] To solve the above technical problems, the present application provides a fault determination method for a micro-service system, comprising: obtaining abnormal monitoring information of the micro-service system, and determining an abnormal event stream according to the abnormal monitoring information; wherein the monitoring information at least includes index information, log information and call chain information; obtaining entity information in the abnormal event stream and the relationship between entities; wherein the entity information in the abnormal event stream is determined by event description information of the abnormal event stream; determining an event causal graph according to the entity information and the relationship between entities; determining an abnormal service in the event causal graph, a propagation path related to the abnormal service and a fault classification of the abnormal service by using a pre-trained model-driven agent in combination with the event causal graph and a pre-set fault mode.

[0007] In an aspect, determining the abnormal event stream according to the abnormal monitoring information comprises: determining an index abnormal event according to the abnormal index information; wherein the index at least comprises a utilization rate index of a processor and a delay index of a system; determining a log abnormal event according to the abnormal log information; determining an abnormal tracking event according to the abnormal call chain information; determining an abnormal event stream according to the index abnormal event, the log abnormal event and the abnormal tracking event.

[0008] In another aspect, the monitoring information is index information, and obtaining the abnormal index information of the microservice system comprises: obtaining an observation value and a preset value of a target index; in a case where a deviation between the observation value and the preset value is greater than a preset deviation, determining that the target index is an abnormal index; obtaining abnormal index information; wherein the abnormal index information at least comprises a time when the abnormal index occurs, a service affected by the abnormal index, a service instance affected by the abnormal index, an abnormal index name, an abnormal type and an abnormal score of the abnormal index; the information of the index abnormal event at least comprises a unique identifier of the event and the abnormal index information.

[0009] In another aspect, determining the log abnormal event according to the abnormal log information comprises: extracting entity information and events in the abnormal log information through a pre-trained model; wherein the entity information in the abnormal log information at least comprises a service, a resource, an error code and a request identifier; determining a type of the extracted event according to a pre-set fault mode or event type; wherein the type of the event at least comprises a connection failure event type and a business logic error event type.

[0010] In another aspect, the monitoring information is call chain information; obtaining the abnormal call chain information of the microservice system comprises: constructing a service topology graph according to call chain data, and obtaining information of the service topology graph; wherein the information of the service topology graph at least comprises a dependency direction between services, a call frequency of a service and a call delay; determining abnormal call chain information according to the information of the service topology graph; the information of the abnormal tracking event at least comprises a type of the event, a tracking identifier, a root service, a bottleneck service and an error type.

[0011] In another aspect, it further comprises: storing the service topology graph and information of the service topology graph in a database; and determining an inter-service invocation event according to the information of the service topology graph; The information of the inter-service invocation event at least includes: event type, target invoker service, source invoker service, and delay.

[0012] On the other hand, obtaining entity information in the abnormal event stream and relationships between entities includes: inputting the abnormal event stream and the service topology graph stored in the database into a pre-trained model; extracting entity information in the abnormal event stream and relationships between entities by the pre-trained model; The entity information in the abnormal event stream at least includes: service name, instance identifier, associated performance indicators, abnormal type, and resources extracted from the event description of the abnormal event stream. The relationships between entities at least include: dependency relationship determined based on the abnormal tracking event, causal relationship determined based on the error stack in the log abnormal event, propagation relationship determined based on the error transmission on the tracking path, and association relationship determined based on the association of different modes of events with request identifier and thread identifier.

[0013] On the other hand, determining the event causal graph according to the entity information and the relationships between entities includes: The entity information is taken as a node of the event causal graph. The relationships between entities are taken as edges between nodes of the event causal graph. After determining the event causal graph according to the entity information and the relationships between entities, further comprising: at least obtaining indicator abnormal value and log error summary; determining attribute information of nodes and attribute information of edges between nodes in the event causal graph based on the indicator abnormal value and the log error summary.

[0014] On the other hand, after determining the event causal graph according to the entity information and the relationships between entities, further comprising: storing the event causal graph, attribute information of nodes and attribute information of edges between nodes in the event causal graph in the database.

[0015] On the other hand, before determining the abnormal service in the event causal graph, the propagation path related to the abnormal service, and the fault classification of the abnormal service by using the pre-trained model driven agent and combining the event causal graph and the pre-set fault mode, further comprising: detecting a fault condition of the microservice system by the agent coordinator; wherein the fault condition includes a fault occurrence condition or a fault non-occurrence condition; when the agent coordinator detects that the fault condition of the microservice system is the fault occurrence condition, sending prompt information for representing fault detection to a target agent; The pre-trained model driven agent determines the abnormal service in the event causal graph, the propagation path related to the abnormal service, and the fault classification of the abnormal service by combining the event causal graph and the pre-set fault mode. The pre-trained model driven agent determines the abnormal service in the event causal graph, the propagation path related to the abnormal service, and the fault classification of the abnormal service by combining the event causal graph and the pre-set fault mode.

[0016] On the other hand, the pre-trained model driven agent determines the abnormal service in the event causal graph, the propagation path related to the abnormal service, and the fault classification of the abnormal service by combining the event causal graph and the pre-set fault mode. In the target agent, the state of the event causal graph is obtained through the pre-trained model; an abnormal signal is identified according to the state of the event causal graph, and a graph query tool is called to extract an abnormal service related to the abnormal signal from the event causal graph; The graph query tool and the knowledge base retrieval tool are called to query the propagation path of the abnormal signal in the microservice system, and a causal chain causing the abnormal service to be abnormal is determined according to the propagation path; The fault classification of the abnormal service is determined according to the fault classification stored in the database.

[0017] On the other hand, the target agent at least includes a first type target agent, a second type target agent and a third type target agent; When the agent coordinator detects that the fault condition of the microservice system is the fault occurrence condition, the prompt information for representing fault detection is sent to the target agent by the agent coordinator; When the agent coordinator detects that the fault condition of the microservice system is the fault occurrence condition, the prompt information for representing starting fault detection is sent to the first type target agent by the agent coordinator; The pre-trained model driven agent determines the abnormal service in the event causal graph, the propagation path related to the abnormal service, and the fault classification of the abnormal service by combining the event causal graph and the pre-set fault mode. In the first type of target agent, the state of the event causal graph is obtained through the pre-trained model; an abnormal signal is identified according to the state of the event causal graph, and a graph query tool is called to extract an abnormal service related to the abnormal signal from the event causal graph; the abnormal service related to the abnormal signal is sent to the agent coordinator; so that the agent coordinator determines a candidate causal chain causing the abnormal service to be abnormal according to the event causal graph, and sends prompt information for representing fault detection of the candidate causal chain to the second type of target agent; In the second type of target agent, a graph query tool and a knowledge base retrieval tool are called to query a propagation path of the abnormal signal in the microservice system, and it is determined that the causal chain causing the abnormal service to be abnormal is the candidate causal chain according to the propagation path; information containing the candidate causal chain is sent to the agent coordinator; so that the agent coordinator sends prompt information for representing fault classification to the third type of target agent; In the third type of target agent, the fault classification of the abnormal service is determined according to the fault classification stored in the database.

[0018] On the other hand, calling the graph query tool to extract the abnormal service related to the abnormal signal from the event causal graph includes: In the first type of target agent, it is judged whether the result returned by the graph query tool meets the first preset requirement; If not, according to the result returned by the graph query tool and the first preset requirement, a thought chain reasoning is performed to generate a new thinking path and determine the query parameters of the result returned by the graph query tool based on the new thinking path; the result returned by the graph query tool according to the query parameters is obtained; it is returned whether the result returned by the graph query tool meets the first preset requirement; If yes, the abnormal service related to the abnormal signal is determined according to the result returned by the graph query tool; Calling the graph query tool and the knowledge base retrieval tool to query the propagation path of the abnormal signal in the microservice system includes: In the second type of target agent, it is judged whether the result returned by the graph query tool and the knowledge base retrieval tool meets the second preset requirement; If not, according to the result returned by the graph query tool and the knowledge base retrieval tool, and the second preset requirement, a thought chain reasoning is performed to generate a new thinking path and determine the query parameters of the result returned by the graph query tool and the knowledge base retrieval tool based on the new thinking path; the result returned by the graph query tool and the knowledge base retrieval tool according to the query parameters is obtained; it is returned whether the result returned by the graph query tool and the knowledge base retrieval tool meets the second preset requirement; If yes, the propagation path of the abnormal signal in the microservice system is determined according to the results returned by the graph query tool and the knowledge base retrieval tool.

[0019] On the other hand, before the intelligent agent coordinator sends the prompt information for characterizing the fault detection of the candidate causal chain to the second type of target intelligent agent, the intelligent agent coordinator further comprises: acquiring the actual fault determination progress and the load condition of the intelligent agent; determining the second type of target intelligent agent according to the actual fault determination progress and the load condition of the intelligent agent; Before the intelligent agent coordinator sends the prompt information for characterizing the fault classification to the third type of target intelligent agent, the intelligent agent coordinator further comprises: determining the third type of target intelligent agent according to the actual fault determination progress and the load condition of the intelligent agent.

[0020] On the other hand, the graph query tool is at least used to perform queries on nodes, edges and paths in the event causal graph; The knowledge base retrieval tool is at least used to query historical fault cases, service metadata, operation and maintenance manuals, fault modes, fault solutions in the database, and provide context and prior knowledge for reasoning.

[0021] On the other hand, the intelligent agent coordinator further comprises: acquiring the thinking path of each intelligent agent, the called tool, the query parameter of the tool and the result returned by the tool; outputting information containing the thinking path of each intelligent agent, the called tool, the query parameter of the tool and the result returned by the tool.

[0022] In order to solve the above technical problems, the present application further provides a fault determination device for a microservice system, comprising: A first acquisition module is configured to acquire abnormal monitoring information of a microservice system and determine an abnormal event stream according to the abnormal monitoring information; wherein the monitoring information at least includes index information, log information and call chain information; A second acquisition module is configured to acquire entity information in the abnormal event stream and the relationship between entities; wherein the entity information in the abnormal event stream is determined by event description information of the abnormal event stream; A first determination module is configured to determine an event causal graph according to the entity information and the relationship between entities; A second determination module is configured to determine an abnormal service in the event causal graph, a propagation path related to the abnormal service and a fault classification of the abnormal service by using a pre-trained model driven intelligent agent and combining the event causal graph and a pre-set fault mode.

[0023] To solve the above technical problems, the present application further provides a computer program product comprising computer programs / instructions which, when executed by a processor, implement the steps of the fault determination method of the microservice system described above.

[0024] To solve the above technical problems, the present application further provides an electronic device comprising: a memory for storing computer programs; a processor for executing the computer programs to implement the steps of the fault determination method of the microservice system described above.

[0025] To solve the above technical problems, the present application further provides a computer readable storage medium having computer programs stored thereon, the computer programs being executed by a processor to implement the steps of the fault determination method of the microservice system described above.

[0026] The present application has the following beneficial effects. Firstly, in the method, after obtaining the abnormal monitoring information of the microservice system, the abnormal event stream is determined according to the abnormal monitoring information; the abnormal event causal diagram is determined based on the abnormal event stream, and then the pre-trained model driven agent is used in combination with the event causal diagram and the fault mode standard to determine the fault point, the fault chain and the fault classification. That is, the method realizes the abnormal detection, fault classification and root cause positioning of the microservice system, and realizes the closed-loop diagnosis from abnormal discovery to root cause positioning. Secondly, compared with the method of determining the microservice fault based on shallow statistical features, in the method provided by the present application, the extracted data is subjected to deep feature engineering and semantic abstraction from the perspective of events and causal relationships, and the system behavior and dependency relationship dispersed in different modal data are fused into a unified event causal diagram through graph modeling technology, providing comprehensive and high-dimensional input for the fault reasoning of the agent, and improving the accuracy of microservice fault determination. Thirdly, when determining the microservice fault, the pre-trained model driven agent is used for analysis, so that the general knowledge and semantic understanding ability learned by the pre-trained model can be relied on for fault analysis, and the efficiency and accuracy of microservice fault determination are improved.

[0027] In addition, the abnormal event stream is determined according to the index abnormal event, the log abnormal event and the abnormal tracking event, breaking the information limitation of single event dimension, so as to more accurately and comprehensively locate the abnormal root cause and improve the accuracy of fault determination.

[0028] When acquiring the abnormal call chain information of the microservice system, first, the service topology graph is constructed according to the call chain data, and then the information of the service topology graph is acquired; finally, the abnormal call chain information is determined according to the information of the service topology graph. By constructing the service topology graph, the user can intuitively understand the call relationship between the services in the microservice system; and based on the information of the service topology graph such as the service interdependence direction, the service call frequency and the call delay, the abnormal call chain information can be determined more accurately.

[0029] By extracting entity information and the relationship between entities in the abnormal event stream through the pre-trained model, the dependence on domain prior knowledge and artificial rule construction can be effectively reduced, and stronger cross-domain adaptability and complex text understanding capability can be possessed, so that the efficiency and accuracy of entity and relationship extraction can be improved while reducing the artificial cost investment, which is especially suitable for processing unstructured and semantically complex text data scenarios.

[0030] By taking the entity information and the relationship between entities as the nodes and edges of the event causal graph respectively, and combining the index abnormal value and the log error summary to supplement the attribute information of the nodes and edges, the event causal graph constructed can not only have a clear entity association structure, but also contain accurate abnormal related attribute data, thereby significantly improving the depth of the event causal graph in describing the event logic and the information integrity.

[0031] The agent coordinator detects the fault condition of the microservice system, sends prompt information of the fault detection to the target agent after detecting the fault, and then uses the target agent driven by the pre-trained model to analyze the fault. That is, the agent coordinator realizes the macroscopic detection of the microservice system fault and the unified management of the target agent.

[0032] Different types of agents (such as the first type of target agent for abnormal signal recognition, the second type of target agent for causal chain determination, and the third type of target agent for fault classification processing) are triggered by the agent coordinator, and the abnormal detection, fault classification and root cause positioning are completed by different types of agents together. This collaborative mechanism effectively overcomes the limitations of traditional single models and fixed expert architectures, improves the accuracy, robustness and explainability of diagnosis, and better adapts to the dynamic changes of the microservice environment.

[0033] A dynamic and adaptive multi-agent collaborative architecture is adopted instead of a fixed expert approach. The agent has the ability of autonomous planning and tool invocation, and can dynamically select the most suitable tool and reasoning path according to the real-time state of the event causal graph and the diagnosis target, and iteratively converge to the root cause. This flexibility enables the system to better adapt to the dynamic changes of the microservice environment and cope with unknown or newly emerging fault modes without the need to frequently retrain the entire model, significantly improving the robustness and generalization ability of the system.

[0034] The output includes the thinking path of each agent, the called tool, the query parameter of the tool, and the result returned by the tool. That is, the system not only gives a diagnostic conclusion (fault point, fault chain, and fault classification), but also records in detail the observation, thinking, and action (including each tool call and its result) of the agent at each step. The operation and maintenance personnel can accurately track each step of logic of the model from the original data to the final decision, greatly improving the explainability, credibility, and auditability of the diagnostic result. This enables the operation and maintenance personnel to trust the judgment of the model, conduct in-depth verification, and learn from it, making up for the shortcomings of traditional Artificial Intelligence (AI) black box models, which is particularly important for production environments with high safety and compliance requirements.

[0035] In addition, the present application also provides a fault determination apparatus of a micro-service system, a computer program product, an electronic device, and a computer readable storage medium, which have the same or corresponding technical features as the above-mentioned fault determination method of a micro-service system, and the same effects. BRIEF DESCRIPTION OF DRAWINGS

[0036] In order to more clearly illustrate the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0037] Figure 1 A flowchart of a fault determination method of a micro-service system provided for the embodiments of the present application; Figure 2 A structural diagram of a large model driven micro-service fault diagnosis system based on adaptive multi-modal agent collaborative reasoning provided for the embodiments of the present application; Figure 3 A structural diagram of an electronic device provided for the embodiments of the present application. DETAILED DESCRIPTION

[0038] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.

[0039] The core of the present application is to provide a fault determination method, apparatus, product, and electronic device of a micro-service system, to solve the problem of low accuracy in analyzing the fault condition of a micro-service system based on the shallow features of the micro-service system.

[0040] In order to make the person skilled in the art better understand the present application, the present application will be further described in detail below in combination with the drawings and specific embodiments. Figure 1 A flowchart of a fault determination method of a microservice system provided by an embodiment of the present application is shown in Figure 1 The method comprises the following steps: S10: acquiring abnormal monitoring information of the microservice system, and determining an abnormal event stream according to the abnormal monitoring information; wherein the monitoring information at least comprises index information, log information and call chain information; S11: acquiring entity information in the abnormal event stream and the relationship between entities; wherein the entity information in the abnormal event stream is determined by event description information of the abnormal event stream; S12: determining an event causal graph according to the entity information and the relationship between entities; S13: utilizing a pre-trained model driven agent, and combining the event causal graph and a pre-set fault mode, to determine an abnormal service in the event causal graph, a propagation path related to the abnormal service and a fault classification of the abnormal service.

[0041] In order to improve the accuracy of fault determination of the microservice system, the abnormal event stream is determined according to the abnormal monitoring information, which comprises: determining an index abnormal event according to abnormal index information; wherein the index at least comprises a utilization rate index of a processor and a delay index of a system; determining a log abnormal event according to abnormal log information; determining an abnormal tracking event according to abnormal call chain information; determining the abnormal event stream according to the index abnormal event, the log abnormal event and the abnormal tracking event.

[0042] Specifically, the monitoring information is index information, and the abnormal index information of the microservice system is acquired, which comprises: acquiring an observation value and a preset value of a target index; in the case where a deviation between the observation value and the preset value is greater than a preset deviation, determining that the target index is an abnormal index; acquiring abnormal index information; wherein the abnormal index information at least comprises a time when the abnormal index occurs, a service affected by the abnormal index, a service instance affected by the abnormal index, an abnormal index name, an abnormal type and an abnormal score of the abnormal index; the information of the index abnormal event at least comprises a unique identifier of the event and the abnormal index information.

[0043] The log abnormal event is determined according to the abnormal log information, which comprises: Extract entity information and events in abnormal log information through a pre-trained model; wherein the entity information in the abnormal log information at least includes services, resources, error codes and request identifiers; Determine the type of the extracted event according to a pre-set fault mode or event type; wherein the type of the event at least includes a connection failure event type and a business logic error event type.

[0044] The monitoring information is call chain information; the abnormal call chain information of the microservice system is obtained by: Construct a service topology graph according to the call chain data, and obtain information of the service topology graph; wherein the information of the service topology graph at least includes the dependency direction between services, the call frequency of the services and the call delay; Determine the abnormal call chain information according to the information of the service topology graph; The information of the abnormal tracking event at least includes: the event type, the tracking identifier, the root service, the bottleneck service and the error type.

[0045] In addition, the inter-service call event can also be determined according to the information of the service topology graph; wherein the information of the inter-service call event at least includes: the event type, the target caller service, the source caller service and the delay.

[0046] In order to obtain the service topology graph, in the implementation, the fault determination method of the microservice system further includes: storing the service topology graph and the information of the service topology graph in a database.

[0047] The above describes the process of obtaining abnormal monitoring information of a microservice system and determining an abnormal event stream according to the abnormal monitoring information; wherein the monitoring information at least includes index information, log information and call chain information. The module that implements this function is called a multi-modal semantic extraction module. In order to enable those skilled in the art to better understand the function implemented by the multi-modal semantic extraction module, the heterogeneous data multi-level semantic extraction mechanism will be described again in conjunction with specific embodiments.

[0048] The present application introduces a more detailed multi-level semantic extraction strategy, aiming to extract structured and semanticized "events" from the original massive heterogeneous monitoring data, and to provide high-quality input for subsequent event causal graph construction and agent reasoning.

[0049] 1) Index data processing: index eventization and anomaly detection.

[0050] Time series prediction and anomaly detection: for each index, identify abnormal points or abnormal segments by calculating the deviation between the observed value and the predicted value (based on the Z-score method).

[0051] Event Conversion: Once an anomaly is detected, it is encapsulated into an "Anomaly Event". Each event contains: a unique identifier (event_id), a timestamp, the affected service (service_name), the affected service instance (instance_id), the anomaly metric name (metric_name, e.g., CPU_Usage, P99_Latency), the anomaly type (anomaly_type, e.g., spike, drop, sustained_high, sudden_change), the anomaly score (anomaly_score, representing the quantified anomaly degree), and other relevant metadata (tags, e.g., deployment_version).

[0052] This process enables the conversion from continuous time-series data to discrete, semanticized "Anomaly Events", greatly simplifying subsequent processing.

[0053] 2) Log Data Processing: Log Structuring and Anomaly Event Extraction.

[0054] Keyword and Regular Expression Matching: Quickly filter out a large number of normal logs.

[0055] LLM-driven Anomaly Event Extraction and Classification: Combined with filtered text, LLM performs information extraction (Information Extraction) to identify key entities (services, resources, error codes, request IDs, etc.) and event types (connection failure, OOM, deadlock, business logic error, etc.). LLM is explicitly prompted to map log content to predefined fault patterns or event types (from general knowledge base). For example, the prompt may include "Based on the following log content, extract key entities and events, and classify them into known fault patterns, or identify as a new anomaly event if no match is found."

[0056] Output: Structured "Log Event Stream". Each event contains: event_id, timestamp, service_name, instance_id.

[0057] semantic_event_type: LLM-identified semantic event type (e.g., DB_Connection_Refused, Application_Crash, API_Auth_Failure).

[0058] original_log_summary: A concise semantic summary of the original log.

[0059] 3) Trace data processing: trace context and service behavior analysis, extract "inter-service call events" and "anomaly trace events" from the original call chain.

[0060] Dynamic topology construction and update: continuously construct and maintain service topology graph G=(V,E) from call chain data, including inter-service dependency direction, call frequency, average delay, etc. This graph will be stored in the general knowledge base for agent query.

[0061] P95 threshold anomaly detection: combined with dynamic threshold, identify response time abnormal Span (call relationship) or entire Trace (call chain).

[0062] Span attribute analysis: extract key attributes of each Span, such as service.name, operation.name, duration, status.code, error label, and associated log event ID.

[0063] Event transformation: encapsulate analysis results as "trace event stream".

[0064] Inter-service call event: {“event_type”:“ServiceCall”,“caller_service”:“UserService”,“callee_service”:“OrderService”,“latency_ms”:50,“success”:true}.

[0065] Anomaly trace event: {“event_type”:“TraceAnomaly”,“trace_id”:“abc”,“root_service”:“Gateway”:“bottleneck_service”:“PaymentService”,“error_type”:“Timeout”}.

[0066] After determining the abnormal event stream, obtain the entity information in the abnormal event stream and the relationship between the entities, including: Input the abnormal event stream and the service topology graph stored in the database into the pre-trained model; Extract entity information and relationships in the abnormal event stream through the pre-trained model; Among them, the entity information in the abnormal event stream at least includes the service name, instance identifier, associated performance indicators, anomaly type and resources extracted from the event description of the abnormal event stream; The relationship between entities includes at least: a dependency relationship determined based on an abnormal tracking event, a causal relationship determined based on an error stack in a log abnormal event, a propagation relationship determined based on an error transmission on a tracking path, and an association relationship determined based on association of different modes of events based on request identification and thread identification.

[0067] Determining the event causal graph according to the entity information and the relationship between entities includes: Taking the entity information as a node of the event causal graph; Taking the relationship between entities as an edge between nodes of the event causal graph; After determining the event causal graph according to the entity information and the relationship between entities, further comprising: At least obtaining an index abnormal value and a log error summary; Determining attribute information of a node and attribute information of an edge between nodes in the event causal graph based on the index abnormal value and the log error summary.

[0068] After determining the event causal graph according to the entity information and the relationship between entities, further comprising: Storing the event causal graph, attribute information of a node in the event causal graph, and attribute information of an edge between nodes in the event causal graph in a database.

[0069] The process of determining the event causal graph is described above. The module that implements this function is called the event causal graph modeling and system state representation module. In order for those skilled in the art to better understand the function implemented by the event causal graph modeling and system state representation module, the following will be described again in conjunction with specific embodiments.

[0070] The core of the event causal graph modeling and system state representation module is to dynamically build and maintain an event causal graph (ECG_t), which represents the behavior, abnormality and causal relationship between them of the microservice system within a specific time window. This graph will be the main "observation object" for the agent to conduct high-level reasoning.

[0071] 1) Event aggregation and time window management: aggregate the index event stream, log event stream and tracking event stream from the previous stage within a specific time window (for example, the past 5 minutes). These events constitute the "original evidence" of ECG_t. Using a sliding time window mechanism, ensure that ECG_t always reflects the latest system state.

[0072] 2) LLM-driven entity-relation extraction and graph construction: this is the core of this module. The present application uses a special LLM as an entity-relation extractor.

[0073] Input: Aggregated event stream (JSON array), and the latest service topology and metadata in the general knowledge base.

[0074] Prompt example: "You are a professional graph knowledge builder. Given the following list of microservice events and system topology, identify the key entities (services, instances, resources, fault types, request IDs, etc.) and their interrelationships (call relationships, exception propagation, causal associations, property associations). Build a graph structure representing these entities and relationships in JSON or GraphML format. For example: 'entity': 'OrderService', 'type': 'Service', 'properties': {'cpu_util': 'high'},'relations': [{'target': 'DB', 'type': 'depends_on','reason': 'DB_Deadlock'}]. Pay special attention to entity ID and type consistency."

[0075] Core logic as follows: (1) Entity recognition: Extract service names, instance IDs, KPIs (Key Performance Indicators), error codes, exception types, resources (such as databases, caches, message queues), etc. from event descriptions as nodes (Nodes) of the graph.

[0076] (2) Relation extraction: Identify the edges (Edges) between nodes.

[0077] (3) Dependency relationship: Based on service calls in trace events.

[0078] (4) Causal relationship: Based on error stacks in log events (e.g., causing a service exception).

[0079] (5) Propagation relationship: Based on error transmission on the trace path.

[0080] (6) Association relationship: Based on request ID, thread ID, and other metadata to associate events of different modalities (e.g., metric exceptions, log errors, and call chain timeouts under a certain request ID).

[0081] (7) Property association: Treat metric exception values, log error summaries, etc. as node or edge properties.

[0082] Graph structure representation: The generated ECG_t is a directed and weighted graph. Both nodes and edges have rich attributes. Node attributes: service nodes can have CPU_anomaly_score, Memory_anomaly_score; error event nodes can have error_type, count; resource nodes can have status. Edge attributes: dependency edges can have avg_latency, error_rate, frequency; causal edges can have causal_strength.

[0083] 3) Graph storage and query optimization: After ECG_t is dynamically constructed, it is stored in the graph database, providing graph query tools for intelligent agents to efficiently retrieve specific patterns, paths, or associated information in the graph. For example, query "all services with CPU anomalies related to databases".

[0084] After determining the event causal graph, the pre-trained model-driven intelligent agent is used in combination with the event causal graph and the pre-set fault mode to determine the abnormal service in the event causal graph, the propagation path related to the abnormal service, and the fault classification of the abnormal service.

[0085] In addition, before using the pre-trained model-driven intelligent agent in combination with the event causal graph and the pre-set fault mode to determine the abnormal service in the event causal graph, the propagation path related to the abnormal service, and the fault classification of the abnormal service, it further includes: detecting the fault condition of the microservice system through the intelligent agent coordinator; wherein the fault condition includes a fault occurrence condition or a fault non-occurrence condition; when the intelligent agent coordinator detects that the fault condition of the microservice system is the fault occurrence condition, sending prompt information representing fault detection to the target intelligent agent; Using the pre-trained model-driven intelligent agent in combination with the event causal graph and the pre-set fault mode to determine the abnormal service in the event causal graph, the propagation path related to the abnormal service, and the fault classification of the abnormal service includes: Using the pre-trained model-driven target intelligent agent in combination with the event causal graph and the pre-set fault mode to determine the abnormal service in the event causal graph, the propagation path related to the abnormal service, and the fault classification of the abnormal service.

[0086] In some embodiments, using the pre-trained model-driven target intelligent agent in combination with the event causal graph and the pre-set fault mode to determine the abnormal service in the event causal graph, the propagation path related to the abnormal service, and the fault classification of the abnormal service includes: In the target agent, the state of the event causal graph is obtained through the pre-trained model; an abnormal signal is identified according to the state of the event causal graph, and a graph query tool is called to extract an abnormal service related to the abnormal signal from the event causal graph; The graph query tool and the knowledge base retrieval tool are called to query the propagation path of the abnormal signal in the microservice system, and a causal chain leading to the abnormality of the abnormal service is determined according to the propagation path; The fault classification of the abnormal service is determined according to the fault classification stored in the database.

[0087] The graph query tool is at least used to perform queries on nodes, edges and paths in the event causal graph; The knowledge base retrieval tool is at least used to query historical fault cases, service metadata, operation and maintenance manuals, fault modes, fault solutions in the database, and to provide context and prior knowledge for reasoning.

[0088] Traditional single model architecture, even based on large pre-trained models, faces problems such as context length limitation, key information loss and inconsistent reasoning path. In the face of massive heterogeneous data, a single model is difficult to effectively integrate all information and maintain the consistency of reasoning. For example, directly inputting all indicators, logs and tracking data into a single large language model may exceed its context window limit and cause information truncation. Therefore, in the present application, fault analysis is performed based on multiple agents. In some embodiments, the target agent includes at least a first type of target agent, a second type of target agent and a third type of target agent.

[0089] When the agent coordinator detects that the fault condition of the microservice system is a fault, the prompt information for representing fault detection is sent to the target agent, including: When the agent coordinator detects that the fault condition of the microservice system is a fault, the prompt information for representing the start of fault detection is sent to the first type of target agent by the agent coordinator; The target agent driven by the pre-trained model determines the abnormal service in the event causal graph, the propagation path related to the abnormal service and the fault classification of the abnormal service in combination with the event causal graph and the pre-set fault mode, including: In the first type of target agent, the state of the event causal graph is obtained through the pre-trained model; an abnormal signal is identified according to the state of the event causal graph, and a graph query tool is called to extract an abnormal service related to the abnormal signal from the event causal graph; the abnormal service related to the abnormal signal is sent to the agent coordinator; so that the agent coordinator determines the candidate causal chain leading to the abnormality of the abnormal service according to the event causal graph, and sends prompt information for representing fault detection on the candidate causal chain to the second type of target agent; In the second type of target agent, the graph query tool and the knowledge base retrieval tool are called to query the propagation path of the abnormal signal in the microservice system, and the causal chain causing the abnormal service exception is determined as a candidate causal chain according to the propagation path; information containing the candidate causal chain is sent to the agent coordinator; so that the agent coordinator sends prompt information for representing fault classification to the third type of target agent; In the third type of target agent, the fault classification of the abnormal service is determined according to the fault classification stored in the database.

[0090] In some embodiments, the calling of the graph query tool to extract the abnormal service related to the abnormal signal from the event causal graph comprises: In the first type of target agent, it is judged whether the result returned by the graph query tool meets the first preset requirement; If not, the thought chain reasoning is performed according to the result returned by the graph query tool and the first preset requirement, to generate a new thinking path and determine the query parameters of the result returned by the graph query tool based on the new thinking path; the result returned by the graph query tool according to the query parameters is obtained; and it is judged whether the result returned by the graph query tool meets the first preset requirement; If yes, the abnormal service related to the abnormal signal is determined according to the result returned by the graph query tool; The calling of the graph query tool and the knowledge base retrieval tool to query the propagation path of the abnormal signal in the microservice system comprises: In the second type of target agent, it is judged whether the result returned by the graph query tool and the knowledge base retrieval tool meets the second preset requirement; If not, the thought chain reasoning is performed according to the result returned by the graph query tool and the knowledge base retrieval tool, and the second preset requirement, to generate a new thinking path and determine the query parameters of the result returned by the graph query tool and the knowledge base retrieval tool based on the new thinking path; the result returned by the graph query tool and the knowledge base retrieval tool according to the query parameters is obtained; and the step of judging whether the result returned by the graph query tool and the knowledge base retrieval tool meets the second preset requirement is returned; If yes, the propagation path of the abnormal signal in the microservice system is determined according to the result returned by the graph query tool and the knowledge base retrieval tool.

[0091] The first preset requirement and the second preset requirement are to achieve a diagnosis target or to meet a stop condition.

[0092] In addition, before the agent coordinator sends prompt information for representing fault detection of the candidate causal chain to the second type of target agent, it further comprises: Obtaining the actual fault determination progress and the load condition of the agent; Determining the second type of target agent according to the actual fault determination progress and the load condition of the agent; The agent coordinator sends prompt information for characterizing fault classification to the third type of target agent before the prompt information is sent. The third type of target agent is determined according to the actual fault determination progress and the load condition of the agent.

[0093] In order to enable the operation and maintenance personnel to accurately track each step of logic of the model from the original data to the final decision, the fault determination method of the micro-service system further includes: Obtaining the thinking path of each agent, the called tool, the query parameter of the tool, and the result returned by the tool; Outputting information containing the thinking path of each agent, the called tool, the query parameter of the tool, and the result returned by the tool.

[0094] The process described above for determining the abnormal service in the event causal graph, the propagation path related to the abnormal service, and the fault classification of the abnormal service by using the pre-trained model driven agent in combination with the event causal graph and the pre-set fault mode can be implemented by an adaptive multi-agent collaborative reasoning module. In order to enable those skilled in the art to better understand the function of the adaptive multi-agent collaborative reasoning module, the process will be described below in combination with specific embodiments.

[0095] The adaptive multi-agent collaborative reasoning module: It aims to achieve dynamic, iterative and interpretable diagnosis of complex faults by building a group of large language model driven agents with autonomous planning, decision making and tool calling capabilities.

[0096] 1) Agent Orchestrator: It plays the role of the "brain" of the entire agent system, responsible for: initial perception and task initialization: receiving the event causal graph (ECG_t) constructed in the previous stage as the initial observation. According to the abnormal mode in the graph and the preset diagnosis target, select and activate the initial perception agent.

[0097] Dynamic planning and scheduling: coordinate the interaction between multiple agents. According to the current diagnosis progress and agent feedback, dynamically determine the next agent that should be activated, and how to allocate tasks and share information.

[0098] Iteration and convergence control: monitor the diagnosis process, including reasoning depth, time limit or confidence threshold, until the convergence condition is reached (for example, finding a high-confidence root cause or being unable to further reason).

[0099] Context management: maintain the "shared memory" or "global context" of the entire diagnosis session, ensuring the consistency of information between agents and the traceability of history.

[0100] 2) LLM-driven Agent Swarm and Dynamic Coordination: The present invention builds a group of agents driven by large language models, which are no longer fixed "expert" divisions, but have the ability of autonomous planning, decision-making and tool calling, and can dynamically play different roles and work together according to task requirements. In the diagnosis process, these agents dynamically play roles and cooperate according to the system state and task requirements. Perception Agent (the first type of target agent described above): mainly responsible for identifying and confirming preliminary abnormal signals, extracting the most prominent and relevant abnormal events and affected entities from ECG_t. Reasoning Agent (the second type of target agent described above): mainly responsible for building causal chains and propagation paths, analyzing abnormal information and iteratively exploring causal relationships, and excluding irrelevant factors. Decision Agent: mainly responsible for integrating the evidence and inferences provided by all perception and reasoning agents, combining fault patterns in the general knowledge base, and finally giving the most likely root cause, fault classification, and generating a structured diagnosis report and detailed reasoning process.

[0101] Each agent has the following core capabilities: Perception: the LLM understands the state of the current observed event causal graph (or its subgraph), identifies preliminary abnormal signals, and extracts the most prominent and relevant abnormal events and affected entities from ECG_t. For example, the perception agent may call the graph query tool to find the node with the highest CPU usage.

[0102] Thinking / Planning: the LLM inside the agent conducts deep reasoning and planning based on the current observations (using the ReAct mode, combined with advanced prompt techniques such as Chain-of-Thought), decides the next analysis direction, the tools to be called and their parameters. This internal thinking enables the agent to autonomously build causal chains and propagation paths, and exclude irrelevant factors. For example, after analyzing the abnormal information provided by the perception agent, the reasoning agent will iteratively call the graph query tool and knowledge base retrieval tool to explore the propagation path of the anomaly in the system, and may find the causal chain of "A service CPU high causes B service response slow".

[0103] Acting / Tool Calling: the agent calls various external tools through predefined interfaces to obtain more information or perform specific analysis. The results of tool calling will be fed back to the agent as new observations, driving the next round of "thinking" and "action".

[0104] Common Toolset: Tools that agents can call include but are not limited to: Graph Query Tool: Perform queries on nodes, edges, paths, patterns in ECG_t, e.g., query_graph(query_type='get_neighbors', node_id='OrderService', relation_type='depends_on').

[0105] KB Retrieval Tool: Query the general knowledge base for historical fault cases, service metadata, operation manuals, fault patterns, common solutions, etc., to provide context and prior knowledge for reasoning, e.g., retrieve_kb_entry(query='deadlock common causes').

[0106] Raw Data Query Tool: Directly query raw logs, metrics, or trace storage (such as Elasticsearch, Prometheus, Jaeger) when necessary to obtain more detailed context to assist verification, e.g., query_raw_log(request_id='xyz').

[0107] 3) Adaptive Iterative Reasoning Process. The diagnosis process is an iterative "observe-think-act loop": Observe: The agent coordinator inputs part or all of ECG_t as the current observation to the agent.

[0108] Think / Plan: The LLM inside the agent performs Chain-of-Thought reasoning based on the observation and the current goal, generating a thinking path and the next action plan (i.e., which tool to call and the tool's parameters).

[0109] Act: The agent executes the plan, calls the corresponding external tool, and takes the tool's output as the new observation.

[0110] Loop: The new observation is fed back to the agent, prompting it to perform the next round of "thinking" and "action", until the diagnosis goal is reached or the stopping condition is met. This iterative process allows the agent to dynamically adjust its reasoning strategy based on real-time feedback.

[0111] Generation and Structuring of Reasoning Chain: To maximize interpretability, this invention enhances the generation and structuring of the reasoning chain, making it include the agent's complete "thought trail." The enhanced reasoning chain structure meticulously records every "observation," "thinking," and "action" of the agent throughout the diagnostic process, forming a traceable and auditable trajectory. The reasoning chain not only includes the final inference but also explicitly records each agent's "thinking" (internal planning), "action" (tools invoked and their parameters), and "observation" (results returned by the tools) at each step. Operations personnel can clearly see when and why the agent invoked which tools, and what information the tools returned to drive the reasoning progress, greatly enhancing transparency and verifiability. The entire reasoning chain demonstrates the complete decision-making process from initial anomaly detection to final root cause localization and recommendations, moving beyond a simple accumulation of conclusions.

[0112] The above describes the entire process of fault determination in microservice systems, namely, the large model-driven microservice fault diagnosis process based on adaptive multimodal intelligent agent collaborative reasoning. In order to enable those skilled in the art to better understand the overall technical solution of the present invention, the following description will continue with reference to the accompanying drawings and specific embodiments. Figure 2 A structural diagram of a large model-driven microservice fault diagnosis system based on adaptive multimodal intelligent agent collaborative reasoning, provided in an embodiment of the present invention, is shown below. Figure 2 As shown, the system includes a raw data layer 1, a multimodal semantic extraction module 2, a general knowledge base 3, an event-cause graph modeling and state representation module 4, a multi-agent collaborative reasoning module 5, and a module 6 for generating structured diagnostic reports. The raw data layer 1 includes metrics, logs, and traces. The multimodal semantic extraction module 2 performs metric event identification, structured log extraction, and trace behavior analysis. The general knowledge base 3 stores service topology, fault cases, and domain knowledge. The event-cause graph modeling and state representation module 4 constructs event-cause graphs, extracts entity-relationships, and dynamically updates service states. The multi-agent collaborative reasoning module 5 includes perception / reasoning / decision agents capable of tool calls and adaptive iteration. The module 6 for generating structured diagnostic reports performs anomaly detection (AD), fault classification (FT), and root cause localization (RCL).

[0113] Specifically, the large model-driven microservice fault diagnosis system based on adaptive multimodal intelligent agent collaborative reasoning provided by this invention includes the following modules: 1) General knowledge base: As an auxiliary capability throughout the process, it provides structured knowledge such as system history failures and operation and maintenance experience for each module and agent to query, thereby enhancing the accuracy and efficiency of reasoning.

[0114] 2) Multi-modal semantic extraction module: process raw metrics, logs, call chain data into higher-level "event streams".

[0115] Metrics: transform from raw time-series data into "metric anomaly events", combined with anomaly detection algorithms.

[0116] Logs: emphasize semantic event extraction after log templating, use LLM to classify and extract events.

[0117] Call chain: emphasize extracting "inter-service call events" and "anomaly call events" from the context, as well as behavior patterns.

[0118] The core goal is to provide rich, semantically labeled events for the next stage of graph modeling.

[0119] The multi-modal semantic extraction module is responsible for multi-level structuring, cleaning and preliminary semantic extraction of massive heterogeneous raw monitoring data. This module designs a customized hierarchical processing mechanism for different data types (metrics, logs, traces), aiming to maximize the preservation of semantic information and fault-related features of the original data, while eliminating redundancy and noise, providing high-quality, semantically rich input for subsequent modules.

[0120] 3) Event causal graph modeling and system state representation module: the core is to build a dynamically updated event causal graph (ECG_t). Through the entity-relation extraction ability of LLM, identify node entities such as services, instances, resources, error types from the event stream of the previous stage, and extract their causal, dependency, propagation, association, etc. relationship as an edge.

[0121] The event causal graph modeling and system state representation module focuses on deep feature engineering and semantic abstraction of the refined data from the perspective of events and causal relationships. This module uses graph modeling technology to integrate system behavior and dependency relationships scattered in different modal data into a unified event causal graph, providing comprehensive, high-dimensional input for agent reasoning.

[0122] 4) Adaptive multi-agent collaborative reasoning module: Introduce Agent Orchestrator, responsible for perception, planning and scheduling, use more general and flexible LLM to drive agent group, including perception agent, reasoning agent and decision agent. These agents have tool calling ability, can dynamically call various external tools to query knowledge base, analyze graph, retrieve raw data, execute specific algorithms, etc. The reasoning process follows the autonomous planning paradigm of Observe-Think-Act-Loop, and the agent decides the next step of analysis and tool calling according to the observation of the event causal graph, and iteratively converges to the root cause. This mechanism significantly improves the adaptability, robustness and ability to handle unknown faults of the system.

[0123] The adaptive multi-agent collaborative reasoning module is composed of large language model driven agents with autonomous planning and tool calling capabilities (rather than fixed experts). These agents dynamically select, schedule and iteratively execute according to system state and task requirements, working together to complete anomaly detection, fault classification and root cause location. This collaborative mechanism effectively overcomes the limitations of traditional single model and fixed expert architecture, improving the accuracy, robustness and explainability of diagnosis, and better adapting to the dynamic microservice environment.

[0124] 5) Module for generating structured diagnostic report: The diagnostic report will contain more detailed reasoning chains, which will clearly record the thinking process of the agent and each tool call and its results, further enhancing transparency and verifiability. At the same time, it will recommend repair measures.

[0125] The microservice system fault determination method provided by the present application realizes: 1. Unified end-to-end intelligent diagnosis flow: The present application integrates the traditional three core tasks of anomaly detection (AD), fault classification (FT) and root cause location (RCL) into an end-to-end, highly automated diagnosis process through a carefully designed agent collaborative reasoning mechanism. This avoids the cumbersome switching between different tools and manual integration of results, greatly reducing the integration complexity and operation overhead of the operation and maintenance process. Operation and maintenance engineers can now obtain a complete diagnostic report from a unified system, from anomaly detection to root cause location, greatly shortening the mean time to recovery (MTTR). According to statistics, the reduction of MTTR is a key indicator of the effectiveness of AIOps systems, and this method is expected to reduce it by more than 30%.

[0126] 2. Deep causal relationship discovery and understanding: Unlike existing methods that only perform shallow fusion of data, the present invention can automatically discover, model, and understand complex causal relationships and fault propagation paths between entities from massive heterogeneous events by constructing an event causal graph (ECG_t) and utilizing intelligent agents for deep reasoning. This capability enables the system to go beyond superficial fault symptoms and directly hit the root cause. For example, it can not only identify CPU high, but also trace back to the specific operation or internal deadlock that caused the CPU high, thus discovering complex fault patterns involving cross-service and cross-modal that are difficult to identify by traditional methods.

[0127] 3. Extremely high transparency and auditable reasoning process: The present invention introduces a structured, detailed, and traceable intelligent agent reasoning trace in the diagnostic results. This means that the system not only gives diagnostic conclusions, but also records in detail the "observation", "thinking", and "action" (including each tool call and its result) of the intelligent agent at each step. The operation and maintenance personnel can accurately track every step of logic from the original data to the final decision of the model, greatly improving the explainability, credibility, and auditability of the diagnostic results. This enables the operation and maintenance personnel to trust the model's judgment, conduct in-depth verification, and learn from it, making up for the shortcomings of traditional AI "black box" models. It is particularly important for production environments with high safety and compliance requirements.

[0128] 4. Adaptability and robustness: The present invention adopts a dynamic and adaptive multi-agent collaborative architecture, rather than a fixed expert. Intelligent agents have the ability of autonomous planning and tool invocation, and can dynamically select the most suitable tools and reasoning paths according to the real-time state of the event causal graph and the diagnostic target, and iteratively converge to the root cause. This flexibility enables the system to better adapt to the dynamic changes of the microservice environment and cope with unknown or newly emerging fault patterns without the need for frequent retraining of the entire model, significantly improving the robustness and generalization ability of the system.

[0129] 5. Low false alarm rate and high precision diagnosis: The present method performs deep correlation and cross-validation of index abnormal events, log semantic events, and tracing behavior events in the event causal graph. Multiple intelligent agents collaborate and verify each other from different perspectives, which can effectively identify and exclude false abnormalities or irrelevant information. This multi-modal, multi-agent interactive verification mechanism significantly reduces the false alarm rate and improves the accuracy of fault diagnosis. In actual production environments, it is expected to achieve higher accuracy than single modal or shallow fusion solutions, while reducing MTTD (Mean Time To Detect) to minutes or even seconds.

[0130] 6. Scalability and field adaptability: The system adopts the architecture of separating agents and tool layers, with excellent scalability. In the future, new LLM agent roles (such as fault repair agents, security audit agents) can be easily introduced, or new professional tools (such as code analysis tools, configuration management tools) can be integrated, to cope with the evolving operation and maintenance needs. This architecture makes the invention not only limited to fault diagnosis, but also extended to a wider AIOps field, embodying strong foresight and frontier.

[0131] 7. Significant reduction in operation and maintenance cost and improvement in efficiency: Through highly intelligent automated fault diagnosis, the invention can greatly reduce the time and effort of operation and maintenance engineers in fault discovery, positioning and analysis. This allows the operation and maintenance team to devote more effort to system optimization, architecture improvement and innovative business support, thereby improving overall operation and maintenance efficiency and team value. In large-scale complex systems, this can save a lot of human resources, significantly reduce operating costs, and reduce service interruptions caused by human errors.

[0132] In the above embodiment, the fault determination method for the microservice system is described in detail, and the invention also provides corresponding embodiments of the fault determination device for the microservice system and electronic equipment. It should be noted that the embodiments of the device part are described from two angles, one is based on the functional module angle, and the other is based on the hardware angle.

[0133] The fault determination device for the microservice system provided by the embodiments of the invention is based on the functional module angle, which includes: The first acquisition module is configured to acquire abnormal monitoring information of the microservice system, and determine an abnormal event stream according to the abnormal monitoring information; wherein the monitoring information at least includes index information, log information and call chain information; The second acquisition module is configured to acquire entity information in the abnormal event stream and the relationship between entities; wherein the entity information in the abnormal event stream is determined by the event description information of the abnormal event stream; The first determination module is configured to determine an event causal diagram according to the entity information and the relationship between entities; The second determination module is configured to determine an abnormal service in the event causal diagram, a propagation path related to the abnormal service and a fault classification of the abnormal service by using a pre-trained model driven agent and combining the event causal diagram and a pre-set fault mode.

[0134] Since the embodiments of the device part correspond to the embodiments of the method part, the embodiments of the device part are described in the description of the embodiments of the method part, which will not be described here.

[0135] Figure 3A structural diagram of an electronic device is provided in the embodiments of the present application. The embodiments are based on a hardware perspective, as shown in the figure, the electronic device comprises: Figure 3 a memory 20, configured to store a computer program; a processor 21, configured to execute the computer program to implement the steps of the fault determination method of the micro-service system mentioned in the above embodiments.

[0136] The processor 21 can include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 21 can be implemented in at least one of a hardware form of a digital signal processor (DSP), a field-programmable gate array (FPGA), and a programmable logic array. The processor 21 can also include a main processor and a coprocessor. The main processor is a processor for processing data in a wake-up state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor 21 can be integrated with a graphics processing unit (GPU) that is responsible for rendering and drawing the content required to be displayed on the display screen. In some embodiments, the processor 21 can also include an artificial intelligence (AI) processor for processing machine learning-related computing operations.

[0137] The memory 20 can include one or more computer-readable storage media, which can be non-transitory. The memory 20 can also include a high-speed random access memory and a non-volatile memory such as one or more disk storage devices, flash storage devices. In the embodiments, the memory 20 is at least used to store the following computer program 201, wherein the computer program is loaded and executed by the processor 21, and can implement the related steps of the fault determination method of the micro-service system disclosed in any of the preceding embodiments. In addition, the resources stored in the memory 20 can also include an operating system 202 and data 203, etc., and the storage mode can be temporary storage or permanent storage. The operating system 202 can include Windows, Unix, Linux, etc. The data 203 can include but is not limited to the data involved in the fault determination method of the micro-service system mentioned above.

[0138] In some embodiments, the electronic device can also include a display screen 22, an input / output interface 23, a communication interface 24, a power supply 25, and a communication bus 26.​

[0139] Those skilled in the art can understand that, Figure 3 The structure shown in the figure does not constitute a limitation on the electronic device, and can include more or fewer components than those shown.

[0140] The electronic device provided by the embodiment of the present application comprises a memory and a processor. When the processor executes the program stored in the memory, the following method can be realized: the fault determination method of the micro-service system, and the effects are the same as above.

[0141] The embodiment of the present application also provides a computer program product comprising computer programs / instructions, which, when executed by a processor, realize the steps of the fault determination method of the micro-service system described above.

[0142] Finally, the present application also provides an embodiment corresponding to a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to realize the steps recorded in the above method embodiment.

[0143] It can be understood that if the method in the above embodiment is realized in the form of a software function unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and executes all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (Read-Only Memory, ROM), a random access memory (Random Access Memory, RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0144] The computer-readable storage medium provided by the present application comprises the fault determination method of the micro-service system mentioned above, and the effects are the same as above.

[0145] The fault determination method, device, product and electronic device of the micro-service system provided by the present application are described in detail above. The embodiments in the specification are described in a progressive manner, and each embodiment mainly describes the differences from other embodiments. The same or similar parts of each embodiment can be referred to. For the device disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the related parts can be referred to the method part. It should be pointed out that, for ordinary skilled in the art, without departing from the principle of the present application, some improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the present application.

[0146] It is also to be noted that, as used in this specification and the appended claims, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise. Furthermore, the terms "comprising," "containing," or any other variation thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements is not necessarily limited to only those elements, but can include other elements not expressly listed or inherent to such process, method, article, or apparatus.

Claims

1. A method for failure determination of a microservice system, characterized in that, The method comprises: obtaining abnormal monitoring information of a microservice system, and determining an abnormal event stream according to the abnormal monitoring information; wherein the monitoring information at least comprises index information, log information and call chain information; obtaining entity information in the abnormal event stream and the relationship between entities; wherein the entity information in the abnormal event stream is determined by event description information of the abnormal event stream; determining an event causal diagram according to the entity information and the relationship between entities; determining an abnormal service in the event causal diagram, a propagation path related to the abnormal service and a fault classification of the abnormal service by using a pre-trained model-driven agent in combination with the event causal diagram and a pre-set fault mode. 2.The method for failure determination of a microservice system according to claim 1, characterized in that, Determining the abnormal event stream according to the abnormal monitoring information comprises: determining an index abnormal event according to abnormal index information; wherein the index at least comprises a utilization rate index of a processor and a delay index of a system; determining a log abnormal event according to abnormal log information; determining an abnormal tracking event according to abnormal call chain information; determining the abnormal event stream according to the index abnormal event, the log abnormal event and the abnormal tracking event.

3. The method of claim 2, wherein, The monitoring information is index information, and obtaining abnormal index information of the microservice system comprises: obtaining an observation value and a preset value of a target index; determining the target index as an abnormal index in a case where a deviation between the observation value and the preset value is greater than a preset deviation; obtaining abnormal index information; wherein the abnormal index information at least comprises a time when the abnormal index occurs, a service affected by the abnormal index, a service instance affected by the abnormal index, an abnormal index name, an abnormal type and an abnormal score of the abnormal index; The information of the index abnormal event at least comprises a unique identifier of the event and the abnormal index information.

4. The method of claim 2, wherein, Determining the log abnormal event according to the abnormal log information comprises: extracting entity information and events in the abnormal log information by using a pre-trained model; wherein the entity information in the abnormal log information at least comprises a service, a resource, an error code and a request identifier; determining a type of the extracted event according to a pre-set fault mode or event type; wherein the type of the event at least comprises a connection failure event type and a business logic error event type.

5. The method of claim 2, wherein, The monitoring information is call chain information; Obtaining abnormal call chain information of the microservice system comprises: constructing a service topology graph according to call chain data, and obtaining information of the service topology graph; wherein the information of the service topology graph at least comprises a dependency direction between services, a calling frequency of a service and a calling delay; determining the abnormal call chain information according to the information of the service topology graph; The information of the abnormal tracking event at least comprises an event type, a tracking identifier, a root service, a bottleneck service and an error type.

6. The method of claim 5, wherein, Further comprising: storing the service topology graph and the information of the service topology graph in a database; and determining an inter-service calling event according to the information of the service topology graph; wherein the information of the inter-service calling event at least comprises an event type, a target calling party service, a source calling party service and a delay.

7. The method of claim 6, wherein, Obtaining entity information in the abnormal event stream and the relationship between entities comprises: inputting the abnormal event stream and the service topology stored in the database into a pre-trained model; extracting entity information and relationships between entities in the abnormal event stream through the pre-trained model; wherein the entity information in the abnormal event stream at least includes service name, instance identifier, associated performance indicators, abnormal type and resources extracted from the event description of the abnormal event stream; the relationships between entities at least include dependency relationships determined based on the abnormal tracking events, causal relationships determined based on error stacks in log abnormal events, propagation relationships determined based on error transmission on the tracking path, and association relationships determined based on request identifiers and thread identifiers to associate events of different modes.

8. The method of claim 7, wherein, determining an event causal graph according to the entity information and the relationships between entities includes: taking the entity information as nodes of the event causal graph; taking the relationships between entities as edges between nodes of the event causal graph; after determining the event causal graph according to the entity information and the relationships between entities, further comprising: at least obtaining indicator abnormal values and log error summaries; determining attribute information of nodes and attribute information of edges between nodes in the event causal graph based on the indicator abnormal values and the log error summaries.

9. The method of claim 8, wherein, after determining the event causal graph according to the entity information and the relationships between entities, further comprising: storing the event causal graph, the attribute information of nodes and the attribute information of edges between nodes in the event causal graph in the database.

10. The method of claim 8, wherein, before determining the abnormal service in the event causal graph, the propagation path related to the abnormal service and the fault classification of the abnormal service by the intelligent agent driven by the pre-trained model in combination with the event causal graph and the pre-set fault mode, further comprising: detecting the fault condition of the microservice system through an intelligent agent coordinator; wherein the fault condition includes a fault occurrence condition or a fault non-occurrence condition; sending prompt information representing fault detection to a target intelligent agent when the intelligent agent coordinator detects that the fault condition of the microservice system is the fault occurrence condition; the intelligent agent driven by the pre-trained model in combination with the event causal graph and the pre-set fault mode to determine the abnormal service in the event causal graph, the propagation path related to the abnormal service and the fault classification of the abnormal service includes: the target intelligent agent driven by the pre-trained model in combination with the event causal graph and the pre-set fault mode to determine the abnormal service in the event causal graph, the propagation path related to the abnormal service and the fault classification of the abnormal service.

11. The method of claim 10, wherein, the target intelligent agent driven by the pre-trained model in combination with the event causal graph and the pre-set fault mode to determine the abnormal service in the event causal graph, the propagation path related to the abnormal service and the fault classification of the abnormal service includes: in the target intelligent agent, obtaining the state of the event causal graph through the pre-trained model; identifying an abnormal signal according to the state of the event causal graph, and calling a graph query tool to extract an abnormal service related to the abnormal signal from the event causal graph; The call graph query tool and the knowledge base retrieval tool are used to query a propagation path of the abnormal signal in the microservice system, and a causal chain causing the abnormal service to be abnormal is determined according to the propagation path; A fault classification of the abnormal service is determined according to a fault classification stored in a database.

12. The method of claim 10, wherein, The target agent at least includes a first type target agent, a second type target agent and a third type target agent; In a case where the agent coordinator detects that the fault condition of the microservice system is a fault, the prompt information for representing fault detection is sent to the target agent, including: In a case where the agent coordinator detects that the fault condition of the microservice system is a fault, the prompt information for representing starting fault detection is sent to the first type target agent by the agent coordinator; The target agent driven by the pre-trained model, in combination with the event causal graph and the pre-set fault mode, determines the abnormal service in the event causal graph, the propagation path related to the abnormal service and the fault classification of the abnormal service, including: In the first type target agent, the state of the event causal graph is obtained through the pre-trained model; the abnormal signal is identified according to the state of the event causal graph, and the abnormal service related to the abnormal signal is extracted from the event causal graph by calling the graph query tool; the abnormal service related to the abnormal signal is sent to the agent coordinator; so that the agent coordinator determines a candidate causal chain causing the abnormal service to be abnormal according to the event causal graph, and sends prompt information for representing fault detection of the candidate causal chain to the second type target agent; In the second type target agent, the call graph query tool and the knowledge base retrieval tool are used to query a propagation path of the abnormal signal in the microservice system, and a causal chain causing the abnormal service to be abnormal is determined according to the propagation path; A fault classification of the abnormal service is determined according to a fault classification stored in a database.

13. The method of claim 12, wherein, The call graph query tool extracts the abnormal service related to the abnormal signal from the event causal graph, including: In the first type target agent, it is judged whether the result returned by the graph query tool meets the first preset requirement; If not, the thought chain reasoning is performed according to the result returned by the graph query tool and the first preset requirement, to generate a new thinking path and determine the query parameters of the graph query tool based on the new thinking path; the result returned by the graph query tool according to the query parameters is obtained; the judgment of whether the result returned by the graph query tool meets the first preset requirement is returned; If yes, the abnormal service related to the abnormal signal is determined according to the result returned by the graph query tool; The call graph query tool and the knowledge base retrieval tool are used to query a propagation path of the abnormal signal in the microservice system, including: In the second type of target agent, it is judged whether the results returned by the graph query tool and the knowledge base retrieval tool meet the second preset requirement; If not, according to the results returned by the graph query tool and the knowledge base retrieval tool, and the second preset requirement, a new thinking path is generated by thinking chain reasoning, and the query parameters of the results returned by the graph query tool and the knowledge base retrieval tool are determined based on the new thinking path; the results returned by the graph query tool and the knowledge base retrieval tool are queried according to the query parameters; the step of judging whether the results returned by the graph query tool and the knowledge base retrieval tool meet the second preset requirement is returned; If yes, according to the results returned by the graph query tool and the knowledge base retrieval tool, the propagation path of the abnormal signal in the microservice system is determined.

14. The method of claim 12, wherein, Before the agent coordinator sends prompt information for representing fault detection of the candidate causal chain to the second type of target agent, it further includes: Obtaining the actual fault determination progress and the load condition of the agent; Determining the second type of target agent according to the actual fault determination progress and the load condition of the agent; Before the agent coordinator sends prompt information for representing fault classification to the third type of target agent, it further includes: Determining the third type of target agent according to the actual fault determination progress and the load condition of the agent.

15. The method of claim 11 or 12, wherein, The graph query tool is at least used for executing the query of nodes, edges and paths in the event causal graph; The knowledge base retrieval tool is at least used for querying historical fault cases, service metadata, operation and maintenance manuals, fault modes and fault solutions in the database, and providing context and prior knowledge for reasoning.

16. The method of claim 13, wherein, It further includes: Obtaining the thinking path of each agent, the called tool, the query parameters of the tool and the results returned by the tool; Outputting information containing the thinking path of each agent, the called tool, the query parameters of the tool and the results returned by the tool.

17. A fault determination apparatus of a microservice system, characterized by, It includes: A first obtaining module is configured to obtain abnormal monitoring information of a microservice system, and determine an abnormal event stream according to the abnormal monitoring information; wherein the monitoring information at least includes index information, log information and call chain information; A second obtaining module is configured to obtain entity information in the abnormal event stream and relationships between entities; wherein the entity information in the abnormal event stream is determined by event description information of the abnormal event stream; A first determining module is configured to determine an event causal graph according to the entity information and the relationships between entities; A second determining module is configured to determine an abnormal service in the event causal graph, a propagation path related to the abnormal service and a fault classification of the abnormal service by using a pre-trained model driven agent and combining the event causal graph and a pre-set fault mode.

18. A computer program product comprising computer programs / instructions, characterized in that, The computer program / instructions are executed by the processor to implement the steps of the fault determination method of the microservice system in any one of claims 1 to 16.

19. An electronic device, comprising: It includes: A memory is configured to store a computer program; A processor is configured to execute the computer program to implement the steps of the fault determination method of the microservice system in any one of claims 1 to 16.

20. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the steps of the fault determination method of the micro-service system according to any one of claims 1 to 16.

Citation Information

Patent Citations

  • Micro-service fault detection method and device, storage medium and computer equipment

    CN115357418A

  • Fault root cause positioning method and device, electronic equipment and readable storage medium

    CN115514627A

  • Micro-service fault diagnosis method and system

    CN115640159A

  • Microservice intelligent operation and maintenance system and method oriented to cloud native and application

    CN117009119A

  • Micro-service fault root cause positioning method

    CN117389779A

Cited By

  • Service fault detection method and system

    CN121478664A

  • Service failure detection method and system

    CN121478664B

  • Fault diagnosis method and device, storage medium and electronic equipment

    CN121585525A

  • Micro-service system defect root cause positioning and SOP automatic generation method based on multi-agent collaboration

    CN121881066A

  • Micro-service fault root cause positioning method, system and equipment

    CN122332173A