Fault positioning method and device, electronic equipment and storage medium

By using a large language model to perform semantic and inference analysis on fault alarm information of microservice systems, the problem of low efficiency in fault location and root cause analysis in microservice architecture is solved, and fast and reliable fault location and root cause tracing are achieved.

CN120979919APending Publication Date: 2025-11-18SHANGHAI ZHONG YUAN NETWORK CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511297481.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-11
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing microservice fault location and root cause analysis techniques are difficult to adapt to complex microservice architectures, resulting in low efficiency in fault location and root cause analysis.

Method used

We use a pre-trained large language model to perform semantic analysis on fault alarm information of microservice system, identify the target service associated with the fault, and perform inference analysis by obtaining the operation-related data of the target service, recursively tracing the fault propagation path until the root cause of the fault is determined.

Benefits of technology

It improves the efficiency and intelligence of fault location and root cause analysis, avoids misjudgment of faults, and ensures the reliability and speed of fault location.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120979919A_ABST
    Figure CN120979919A_ABST
Patent Text Reader

Abstract

The invention relates to a fault positioning method and device, electronic equipment and a storage medium. The method comprises the following steps: receiving fault alarm information sent by a micro-service system; performing semantic analysis on the fault alarm information by using a pre-trained large language model based on a topological structure of the micro-service system, and determining a target service associated with the fault; obtaining operation related data of the target service, and performing reasoning analysis on the operation related data of the target service by using a pre-trained second large language model to obtain a fault reason of the target service; and when it is determined that the fault cause of the target service is caused by the upstream service fault of the target service, performing recursive analysis until it is determined that the upstream service causing the fault does not exist, and obtaining the fault root cause of the fault. According to the method, reasoning analysis is carried out through linkage of various operation related data, a fault propagation path is automatically tracked based on a large language model, the efficiency and the intelligent degree of fault positioning and fault root cause analysis can be improved, and the reliability of fault positioning is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of fault location technology, and in particular to a fault location method, apparatus, electronic device and storage medium. Background Technology

[0002] With the development of the Internet, more and more Internet companies are adopting microservice architecture to build highly available distributed systems. These systems are highly dynamic and heterogeneous, supporting various Internet business scenarios. However, if any microservice in a microservice architecture fails, the entire system will crash, affecting user experience. Therefore, when a highly available distributed system built with a microservice architecture fails, microservice fault location technology plays a particularly important role.

[0003] Existing microservice fault location and root cause analysis technologies suffer from weak correlation of alarm information, requiring a significant amount of time and manpower to find and locate the root cause of faults. They are unable to quickly locate complex fault propagation paths, are difficult to adapt to complex microservice architectures, and have low efficiency in fault location and root cause analysis. Summary of the Invention

[0004] This application provides a fault location method, apparatus, electronic device, and storage medium to solve the technical problems that existing microservice fault location and root cause analysis technologies are difficult to adapt to complex microservice architectures and have low efficiency in fault location and root cause analysis.

[0005] In a first aspect, this application provides a fault location method, the method comprising:

[0006] Receive fault alarm information from the microservice system;

[0007] Using a pre-trained first language model, based on the topology of the microservice system, semantic analysis is performed on the fault alarm information to determine the target service associated with the fault.

[0008] Obtain the operation-related data of the target service;

[0009] The second pre-trained language model is used to perform inference analysis on the operation-related data of the target service to obtain the cause of the target service's failure.

[0010] If it is determined that the failure of the target service is caused by a failure of the upstream service of the target service, the upstream service is updated to a new target service, and the process of obtaining the operation-related data of the target service and subsequent steps is repeated until it is determined that there is no upstream service causing the failure, thus obtaining the root cause of the failure.

[0011] In one possible implementation, the step of using a pre-trained first large language model to perform semantic analysis on the fault alarm information based on the topology of the microservice system to determine the target service associated with the fault includes:

[0012] The fault alarm information and the topology of the microservice system are input into a pre-trained first language model, so that the first language model can perform semantic analysis on the fault alarm information based on the topology of the microservice system to determine the target service associated with the fault.

[0013] In one possible implementation, obtaining the operation-related data of the target service includes:

[0014] Call at least one of the distributed tracing interface, service performance metric interface, and log data interface of the target service to obtain the operation-related data of the target service; wherein, the operation-related data includes at least one of the following: distributed tracing data, performance monitoring data, and log monitoring data.

[0015] In one possible implementation, determining that the failure of the target service was caused by a failure of an upstream service of the target service includes:

[0016] Based on the topology of the microservice system, a set of upstream services that have a direct calling relationship with the target service is determined;

[0017] The abnormal features in the operation-related data of each upstream service in the upstream service set are extracted by the pre-trained third language model, and the abnormal features are matched with the failure cause of the target service by time sequence and causal correlation degree calculation.

[0018] If the anomaly of any upstream service occurs earlier than the time of the fault cause of the target service, and the causal correlation exceeds a preset threshold, then the fault cause of the target service is determined to be caused by the fault of that upstream service.

[0019] In one possible implementation, the method further includes:

[0020] If it is determined that there is no upstream service causing the failure for the initially identified target service, candidate upstream services that meet the set conditions are selected from the set of first-level upstream services that have a direct calling relationship with the initially identified target service.

[0021] Obtain the operation-related data of the candidate upstream services;

[0022] The second language model is used to perform reasoning analysis on the operation-related data of the candidate upstream service to determine whether the candidate upstream service is faulty.

[0023] If it is determined that the candidate upstream service is faulty, the candidate upstream service is updated to a new target service, and the process of obtaining the operation-related data of the target service and subsequent steps is returned until it is determined that there is no upstream service causing the fault, and the root cause of the fault is obtained.

[0024] If it is determined that the candidate upstream service is not faulty, the fault cause of the target service initially identified is determined as the root cause of the fault.

[0025] In one possible implementation, the method further includes:

[0026] After obtaining the root cause of the fault, a fault analysis report is generated based on the root cause.

[0027] In one possible implementation, generating a fault analysis report based on the root cause of the fault includes:

[0028] Based on the root cause of the fault and the fault cause of each target service determined in the entire fault localization process, a fault propagation path is generated, and a pre-trained third language model is used to perform reasoning analysis on the root cause of the fault and the fault propagation path to obtain a repair suggestion for the root cause of the fault.

[0029] The fault propagation path, the root cause of the fault, and the repair suggestions for the root cause are integrated into a fault analysis report in a preset format.

[0030] Secondly, this application provides a fault location device, the device comprising:

[0031] The alarm information receiving module is used to receive fault alarm information sent by the microservice system;

[0032] The target service determination module is used to perform semantic analysis on the fault alarm information based on the topology of the microservice system using a pre-trained first language model to determine the target service associated with the fault.

[0033] The data acquisition module is used to acquire operation-related data of the target service;

[0034] The fault cause acquisition module is used to perform reasoning analysis on the operation-related data of the target service using a pre-trained second language model to obtain the fault cause of the target service.

[0035] The fault root cause determination module is used to update the upstream service to a new target service when it is determined that the fault of the target service is caused by the fault of the upstream service of the target service, and return to execute the steps of obtaining the operation-related data of the target service and thereafter, until it is determined that there is no upstream service causing the fault, and thus obtain the fault root cause.

[0036] In one possible implementation, the target service determination module is specifically used for:

[0037] The fault alarm information and the topology of the microservice system are input into a pre-trained first language model, so that the first language model can perform semantic analysis on the fault alarm information based on the topology of the microservice system to determine the target service associated with the fault.

[0038] In one possible implementation, the data acquisition module is specifically used for:

[0039] Call at least one of the distributed tracing interface, service performance metric interface, and log data interface of the target service to obtain the operation-related data of the target service; wherein, the operation-related data includes at least one of the following: distributed tracing data, performance monitoring data, and log monitoring data.

[0040] In one possible implementation, determining that the failure of the target service was caused by a failure of an upstream service of the target service includes:

[0041] Based on the topology of the microservice system, a set of upstream services that have a direct calling relationship with the target service is determined;

[0042] The abnormal features in the operation-related data of each upstream service in the upstream service set are extracted by the pre-trained third language model, and the abnormal features are matched with the failure cause of the target service by time sequence and causal correlation degree calculation.

[0043] If the anomaly of any upstream service occurs earlier than the time of the fault cause of the target service, and the causal correlation exceeds a preset threshold, then the fault cause of the target service is determined to be caused by the fault of that upstream service.

[0044] In one possible implementation, the device is further used for:

[0045] If it is determined that there is no upstream service causing the failure for the initially identified target service, candidate upstream services that meet the set conditions are selected from the set of first-level upstream services that have a direct calling relationship with the initially identified target service.

[0046] Obtain the operation-related data of the candidate upstream services;

[0047] The second language model is used to perform reasoning analysis on the operation-related data of the candidate upstream service to determine whether the candidate upstream service is faulty.

[0048] If it is determined that the candidate upstream service is faulty, the candidate upstream service is updated to a new target service, and the process of obtaining the operation-related data of the target service and subsequent steps is returned until it is determined that there is no upstream service causing the fault, and the root cause of the fault is obtained.

[0049] If it is determined that the candidate upstream service is not faulty, the fault cause of the target service initially identified is determined as the root cause of the fault.

[0050] In one possible implementation, the device further includes:

[0051] The fault analysis report generation module is used to generate a fault analysis report based on the root cause of the fault after obtaining the root cause of the fault.

[0052] In one possible implementation, the fault analysis report generation module is specifically used for:

[0053] Based on the root cause of the fault and the fault cause of each target service determined in the entire fault localization process, a fault propagation path is generated, and a pre-trained third language model is used to perform reasoning analysis on the root cause of the fault and the fault propagation path to obtain a repair suggestion for the root cause of the fault.

[0054] The fault propagation path, the root cause of the fault, and the repair suggestions for the root cause are integrated into a fault analysis report in a preset format.

[0055] Thirdly, this application provides an electronic device, including: a processor and a memory, wherein the processor is configured to execute a fault location program stored in the memory to implement the fault location method described in any one of the first aspects.

[0056] Fourthly, this application provides a storage medium storing one or more programs that can be executed by one or more processors to implement the fault location method described in any one aspect.

[0057] Compared with the prior art, the technical solution provided in this application has the following advantages: The method provided in this application receives fault alarm information issued by a microservice system; utilizes a pre-trained large language model to perform semantic analysis on the fault alarm information based on the topology of the microservice system to determine the target service associated with the fault; obtains the operation-related data of the target service, and uses a pre-trained second large language model to perform reasoning analysis on the operation-related data of the target service to obtain the cause of the fault in the target service; when it is determined that the cause of the fault in the target service is caused by the fault of the upstream service of the target service, recursive analysis is performed until it is determined that there is no upstream service causing the fault, thus obtaining the root cause of the fault. This method automatically performs semantic analysis on the alarm information based on a large language model to obtain the target service causing the fault, and performs reasoning analysis by linking multiple operation data of the target service to avoid misjudgment of faults and ensure the reliability of fault location. Furthermore, when it is determined that the fault of the target service is caused by the upstream service, recursive analysis is performed to automatically track the fault propagation path and locate the root cause of the fault, which improves the efficiency and intelligence of fault location and root cause analysis. Attached Figure Description

[0058] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0059] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0060] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.

[0061] Figure 1 A flowchart illustrating an embodiment of a fault location method provided in this application;

[0062] Figure 2 A flowchart illustrating another embodiment of the fault location method provided in this application;

[0063] Figure 3 A flowchart illustrating another embodiment of the fault location method provided in this application;

[0064] Figure 4 A block diagram illustrating an embodiment of a fault location device provided in this application;

[0065] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0066] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0067] The following disclosure provides numerous different embodiments or examples for implementing various structures of this application. To simplify the disclosure, specific examples of components and arrangements are described below. These are merely examples and are not intended to limit the scope of this application. Furthermore, reference numerals and / or letters may be repeated in different examples. Such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.

[0068] To address the limitations of existing microservice fault location and root cause analysis techniques in adapting to complex microservice architectures and their low efficiency, this application provides a fault location method, apparatus, electronic device, and storage medium. Based on a large language model, it automatically performs semantic analysis on alarm information to identify the target service causing the fault. It then uses various operational data from the target service for inference analysis, avoiding misjudgments and improving the reliability of fault location. When it is determined that the fault in the target service is caused by an upstream service, recursive analysis is performed to automatically trace the fault propagation path and locate the root cause, thereby improving the efficiency and intelligence of fault location and root cause analysis.

[0069] Figure 1 A flowchart illustrating an embodiment of a fault location method provided in this application is shown below. Figure 1 As shown, it includes the following steps:

[0070] Step 101: Receive fault alarm information from the microservice system.

[0071] A microservice system refers to a distributed business support system built by modern internet enterprises based on a microservice architecture to support specific business scenarios. Essentially, it consists of multiple relatively independent microservice units, each focusing on a single business function, such as user authentication services, order processing services, and content storage services. It is inherently highly dynamic and heterogeneous, supporting elastic scaling; for example, automatically adding service instances during peak business periods and automatically reducing instances during off-peak periods. Furthermore, each functional module can be developed based on different technology stacks to meet the technical needs of different business scenarios.

[0072] Because of the high dynamism and heterogeneity of microservice systems, which serve as the underlying support for business operations, any failure will lead to business interruption. If the root cause of the failure cannot be quickly located and remedial measures taken, the entire microservice system will fail, degrading the user experience. Therefore, the solution provided in this application can quickly and intelligently locate the cause and root cause of microservice system failures based on the characteristics of large language models.

[0073] Fault alarm information refers to status notification data triggered and sent by the supporting monitoring and alarm components during the operation of a microservice system, indicating that a certain service or component is malfunctioning. Essentially, it serves as entry point information for fault localization, providing an initial basis for subsequent steps and preventing troubleshooting from becoming aimless and scopeless, thus avoiding prolonged service interruptions.

[0074] In one embodiment, the specific implementation of receiving fault alarm information issued by the microservice system is as follows: when the microservice system fails, the monitoring tool is invoked to extract the structured fault alarm information from the microservice system logs.

[0075] For example, suppose at 8:15 PM one evening, a large number of users reported that their video uploads failed. At this time, the microservice system triggers a fault alarm and calls the microservice system's supporting monitoring tools, such as Prometheus + SkyWalking, to find the running data of the user content upload service and extract the fault alarm information from the running logs.

[0076] Step 102: Using the pre-trained first language model, based on the topology of the microservice system, perform semantic analysis on the fault alarm information to determine the target service associated with the fault.

[0077] The microservice topology can be used to describe the calling relationships, hierarchical structure, and interaction rules among service units (including dependent components) in a microservice system, providing a visual or structured model and serving as the core basis for understanding inter-service dependency logic. In this embodiment, the microservice topology includes at least: the calling dependencies between microservices, the association between microservices and the gateway, and the functional labels of each microservice and the responsibility labels of the gateway. Its representation can be structured data or a visual topology diagram.

[0078] The target service associated with the fault can be a core service unit directly related to the current alarm information. There are two possible relationships between the target service and the fault alarm information: direct and potential. Direct relationship: the target service is explicitly mentioned in the alarm information, for example, if the alarm information mentions a sudden increase in the error rate of the order service, then the order service is the initial target service. Potential relationship: based on the topology, the target service is both the alarm service and a core dependency. For example, if the alarm information mentions an order service timeout, and the topology shows that the order service strongly depends on the inventory service, then the inventory service might be identified as a potential target service.

[0079] In one embodiment, the specific implementation method of using a pre-trained first large language model to perform semantic analysis on fault alarm information based on the topology of the microservice system and determine the target service associated with the fault is as follows: the fault alarm information and the microservice topology graph are standardized, prompt words are generated based on the two types of data after standardization, and the prompt words are input into the pre-trained first large language model so that the model outputs the target service in a preset format.

[0080] For example, based on the Large Language Model Context Protocol, fault alarm information and microservice topology diagrams are standardized to generate structured text prompts, such as: Fault alarm information: "The user-content-upload-service service has the following faults: upload interface failure rate 18%, 90% return 503 errors, key log information: continuous timeouts in calls to the / storage / upload interface of object-storage-service, error type: interface timeout, service unavailable." Microservice topology diagram: Service node: [{"service_name":"user-content-upload-service",}]

[0081] Call relationship: [{"source":"user-content-upload-service","target":"object-storage-service","dependency":3,}].... After receiving the input alarm information and microservice topology diagram, the large language model identifies service entities (such as user-uploaded content services) and fault events (such as interface timeouts) from the alarm information. It parses the service call direction (such as A depending on B) and dependency strength (A strongly depends on B) from the topology structure. After performing reasoning analysis based on the above information, the model outputs a list of target services containing service names, association types, and association criteria. Services that are potentially associated with the alarm information and are strongly dependent on the target services are identified as target services.

[0082] Step 103: Obtain the operation-related data of the target service.

[0083] Operational data refers to the multi-dimensional technical indicators and behavioral records generated by the target service during its operation, which can be used to analyze the causes of failures. Essentially, it is the core basis for determining whether a service is abnormal and locating specific fault points.

[0084] In one embodiment, the specific implementation method for obtaining the operation-related data of the target service is as follows: call the operation monitoring data of the service within a certain time range according to the service name of the target service, and clean and encapsulate the operation monitoring data to obtain the operation-related data.

[0085] For example, a specific time range is determined based on the fault alarm time (e.g., September 5th, 8:15). For instance, the complete time range of the fault occurrence, from 5 minutes before the alarm to 5 minutes after the alarm, is defined as 8:10-8:20. Then, the interfaces of various running monitoring systems (e.g., distributed tracing data interface, performance indicator data interface, and critical log data interface) are called to collect monitoring data within the 8:10-8:20 time range, according to the target service name: user-uploaded content service. The collected data is then standardized and finally encapsulated into structured runtime-related data.

[0086] Step 104: Use the pre-trained second language model to perform reasoning analysis on the operation-related data of the target service to obtain the cause of the target service failure.

[0087] In one embodiment, the specific implementation of using a pre-trained second language model to perform reasoning analysis on the operation-related data of the target service to obtain the cause of the target service failure is as follows: convert the operation-related data of the target service into prompt words that the second language model can understand, input the prompt words into the pre-trained second language model, and clearly define the analysis target as outputting the cause of the target task failure.

[0088] For example, prompts are generated according to a preset prompt template, such as: Task: Analyze the runtime-related data of the target service to locate the cause of the failure. Input data: Runtime-related data of the target service. Analysis requirements: Output format is a JSON array. The prompt template is filled with the obtained runtime-related data to obtain prompts. The prompts are then input into the second language model to obtain the output results.

[0089] Step 105: Determine whether the failure of the target service is caused by a failure of its upstream service. If it is determined that the failure of the target service is caused by a failure of its upstream service, proceed to step 106; if it is determined that there is no upstream service causing the failure, proceed to step 107.

[0090] Step 106: Update the upstream service to the new target service and return to step 103.

[0091] Step 107: Obtain the root cause of the fault.

[0092] The following is a unified description of steps 105-107 above:

[0093] An upstream service can refer to a preceding service unit in a microservice topology that is directly called or depends on by the current target service. For example, in a call relationship where A depends on B, B is A's upstream service. Furthermore, due to the complexity of microservice systems, B may also depend on C, in which case C is B's upstream service. This is merely an illustrative example and is not intended to be limiting.

[0094] The root cause of a failure can refer to the initial source that triggers the entire failure chain, that is, the fundamental reason that exists independently of the failure of other services. For example, if a disk storage service fails to write due to a full disk, which in turn causes an object storage service timeout, which in turn causes the user content upload service to fail, then the root cause of the failure of the user content upload service is that the disk is full.

[0095] In one embodiment, if it is determined that the failure of the target service is caused by a failure of the upstream service of the target service, the upstream service is updated to a new target service, and the process returns to obtain the operation-related data of the target service and the subsequent steps until it is determined that there is no upstream service causing the failure. The specific implementation of obtaining the root cause of the failure is as follows: extract the key field of whether the upstream service is called from the failure cause of the target service. If it is determined that there is a field of calling the upstream service, it is determined that the failure is caused by the upstream service. Obtain the operation-related data of the upstream service, and use the pre-trained second large-scale language model to perform reasoning analysis on the operation-related data of the upstream service to obtain the failure cause until it is determined that there is no upstream service causing the failure. The final failure cause is determined as the root cause of the failure.

[0096] For example, assuming the failure reason for the user content upload service includes a "timeout field for calling upstream services (such as object storage services)," it is determined that there is an upstream service call. The object storage service is then identified as the target service. The monitoring system interface is called to obtain the operation-related status data of the object storage service, and this data is output to the second language model to obtain the failure reason. In the failure reason for the object storage service, there are no fields related to calling upstream service interfaces. Only the field "Due to the disk IO utilization rate remaining above 95% for a long time (threshold 80%), the write request processing delay (average 3200ms) exceeds the 3000ms timeout configuration of the downstream service, causing the user content upload service to fail" is included. This field is then identified as the root cause of the failure.

[0097] The method provided in this application receives fault alarm information from a microservice system; utilizes a pre-trained large language model to perform semantic analysis on the fault alarm information based on the microservice system's topology to determine the target service associated with the fault; acquires the target service's operational data, and uses a pre-trained second large language model to perform inference analysis on the target service's operational data to obtain the cause of the target service's fault; if it is determined that the target service's fault is caused by a fault in an upstream service, recursive analysis is performed until it is determined that there is no upstream service causing the fault, thus obtaining the root cause of the fault. This approach automatically performs semantic analysis on alarm information based on a large language model to obtain the target service causing the fault, and performs inference analysis by linking multiple operational data of the target service, avoiding misjudgment of faults and ensuring the reliability of fault location. Furthermore, when it is determined that the target service's fault is caused by an upstream service, recursive analysis is performed to automatically trace the fault propagation path and locate the root cause of the fault, improving the efficiency and intelligence of fault location and root cause analysis.

[0098] Figure 2 A flowchart illustrating another embodiment of the fault location method provided in this application is shown below. Figure 1Based on the illustrated process, this section primarily describes how to determine if the target service's failure was caused by an upstream service. (See [link to relevant documentation]). Figure 2 As shown, it includes the following steps:

[0099] Step 201: Receive fault alarm information from the microservice system.

[0100] For step 201, please refer to the detailed description of the relevant embodiments above.

[0101] Step 202: Input the fault alarm information and the topology of the microservice system into the pre-trained first language model, so that the first language model can perform semantic analysis on the fault alarm information based on the topology of the microservice system and determine the target service associated with the fault.

[0102] The target service associated with the fault can refer to the core service unit that is directly or potentially associated with the current fault alarm information or the current target service when a fault occurs in the microservice system. It is the key object to be analyzed during the fault localization process.

[0103] In one embodiment, fault alarm information and the topology of the microservice system are input into a pre-trained first language model. The first language model performs semantic analysis on the fault alarm information based on the topology of the microservice system to determine the specific implementation of the target service associated with the fault. The specific implementation method is as follows: a structured prompt word is generated based on the fault alarm information and the microservice topology diagram. The prompt word is input into the pre-trained first language model so that the model combines the logical relationship between various services in the microservice topology diagram to output the target service associated with the fault.

[0104] Step 203: Call at least one of the target service's distributed tracing interface, service performance metrics interface, and log data interface to obtain the target service's operation-related data; wherein, the operation-related data includes at least one of the following: distributed tracing data, performance monitoring data, and log monitoring data.

[0105] A distributed tracing interface can be a standardized interface in a microservice system used to query details of cross-service call chains involving a target service. Its core function is to obtain the flow path of a request across multiple services, the time consumed at each stage, and the status.

[0106] Service performance metrics interfaces refer to standardized interfaces used to query quantitative metrics data during the operation of a target service. Their core function is to obtain metrics data such as service business performance, resource consumption, and dependency calls. These metrics data are used to reflect the health status of the service operation.

[0107] Log data interfaces can refer to standardized interfaces used to query text records output during the operation of a target service. Their core function is to obtain service error logs, business logs, and system logs, and extract details of key events when a failure occurs.

[0108] In one embodiment, at least one of the distributed tracing interface, service performance index interface, and log data interface of the target service is invoked to obtain the operation-related data of the target service; wherein, the operation-related data includes at least one of the following: the specific implementation of distributed tracing data, performance monitoring data, and log monitoring data is as follows: the interface invocation strategy is determined according to the service characteristics of the target service, the interface invocation parameters are constructed according to the interface invocation strategy and the interface invocation operation is executed, and the raw data returned by the interface is integrated and processed to obtain structured operation-related data.

[0109] In addition, an API call protection mechanism can be set up. Specifically, a maximum amount of data to be returned can be set for each API to prevent performance degradation due to massive amounts of data. If an API call fails, the API will be re-called until the set threshold is met or the API call succeeds, ensuring the integrity and reliability of the relevant data.

[0110] Step 204: Use the pre-trained second language model to perform reasoning analysis on the operation-related data of the target service to obtain the cause of the target service failure.

[0111] For step 204, please refer to the detailed description of the relevant embodiments above.

[0112] Step 205: Determine whether the failure of the target service is caused by a failure of the upstream service of the target service. If it is determined that the failure of the target service is caused by a failure of the upstream service of the target service, proceed to step 206; if it is determined that there is no upstream service causing the failure, proceed to step 207.

[0113] Step 206: Update the upstream service to the new target service and return to step 203.

[0114] Step 207: Obtain the root cause of the fault.

[0115] The following is a unified description of steps 205-207 above:

[0116] In one embodiment, the cause of the target service failure is determined to be caused by the failure of an upstream service, based on the topology of the microservice system: a set of upstream services that have a direct calling relationship with the target service is identified; abnormal features in the operation-related data of each upstream service in the upstream service set are extracted using a pre-trained third language model, and the abnormal features are matched with the cause of the target service failure in terms of time sequence and causal correlation are calculated; if the abnormal feature of any upstream service occurs earlier than the cause of the target service failure, and the causal correlation exceeds a preset threshold, then the cause of the target service failure is determined to be caused by the failure of that upstream service.

[0117] The aforementioned abnormal service characteristics include at least abnormal data and specific timestamps, which are essentially key feature data used to reflect the current service abnormality.

[0118] For example, if the target service is determined to be a user-uploaded content service, a pre-trained third-level language model is used to filter upstream services that have direct call relationships with the target service based on the microservice topology graph, generating a set of upstream services, such as a video request processing service and a storage service. Then, runtime-related data for two upstream services in the set is obtained, and abnormal features are extracted from this data (e.g., the video request processing service exhibited abnormal performance metrics at 8:15:00 on September 5th, an abnormal response time at 8:15:20, and an abnormal log error at 8:15:40). Based on preset time-series matching rules, it is determined whether the abnormal occurrence time of the upstream service is earlier than that of the target service. For example, if the earliest abnormal occurrence time for the user-uploaded content service is 8:20:10, while the earliest abnormal occurrence time for the video request service is 8:15:00, the video request service's abnormal occurrence time is earlier than the user-uploaded content service's. Furthermore, if the abnormal features of the video request service completely match those of the user-uploaded content service, and the causal correlation between the two abnormal features identified by the large language model is 90%, exceeding the preset threshold of 60%, then the failure of the target service is determined to be caused by a failure in this upstream service.

[0119] pass Figure 2 The illustrated embodiment describes the following: Based on a large language model and microservice topology graph, fault alarm information is analyzed to avoid the subjectivity and omission risks of manual selection of target services. It enables in-depth analysis of fault alarm information, using a large model to recursively analyze and find the upstream service causing the target service failure, ultimately identifying the root cause. It can penetrate multiple service dependencies to find the root cause, adapting to complex microservice structures. Simultaneously, it intelligently realizes fault location and root cause analysis, improving the efficiency of fault location and enabling timely implementation of corresponding measures to resolve faults, preventing prolonged microservice system failures and enhancing the user experience.

[0120] This application also provides a misjudgment mechanism to avoid misjudging faults. In one embodiment, the specific implementation is as follows: if it is determined that there is no upstream service causing the fault for the initially determined target service, candidate upstream services that meet the set conditions are selected from the set of first-level upstream services that have a direct calling relationship with the initially determined target service; the operation-related data of the candidate upstream services are obtained; and the operation-related data of the candidate upstream services are inferred and analyzed using the second language model to determine whether the candidate upstream services have faults.

[0121] The aforementioned conditions can refer to the upstream service that is called most frequently by the target service. This is just an example, and there may be other conditions. This application does not limit this.

[0122] Candidate upstream services can be upstream service units that need further verification to determine whether they have faults, selected from the set of first-level upstream services that have a direct calling relationship with the "target service that is initially determined to have no upstream faults" in the fault location misjudgment mechanism, based on preset screening conditions (such as calling frequency, dependency strength, etc., which are not limited in this application embodiment).

[0123] For example, assuming that the object storage service is initially determined to have no upstream service causing the failure, in order to avoid misjudgment of the failure, the set of upstream services of the object storage service is obtained from the microservice topology, and upstream services that meet the preset rules are selected from the set of upstream services. For example, if the operation-related data of the target service shows that it calls the backup service most frequently, the backup service is identified as a candidate upstream service, and the operation-related status data of the backup service is obtained. The operation status data is input into the second language model for reasoning analysis to determine whether the backup service has failed.

[0124] In one embodiment, if it is determined that a candidate upstream service is faulty, the candidate upstream service is updated to a new target service, and the process returns to obtain the operation-related data of the target service and subsequent steps until it is determined that there is no upstream service causing the fault, thus obtaining the root cause of the fault.

[0125] For example, assuming the output of the second largest language model is that the backup service has failed, the backup service is updated to a new target service, and the operation-related data of the backup service is obtained. Based on the large language model, the cause of the backup service failure is determined. If the backup service failure is caused by an upstream service, recursive analysis is performed until it is determined that there is no upstream service that caused the failure, and the root cause of the failure is determined.

[0126] In one embodiment, if it is determined that there is no fault in the candidate upstream service, the fault cause of the target service initially identified is determined as the root cause of the fault.

[0127] For example, if the backup service is not faulty, the target service, i.e., the object storage service, is identified as the root cause of the fault.

[0128] Based on the descriptions in the above embodiments, the misjudgment mechanism provided in this application filters candidate upstream services and verifies whether the candidate upstream services are faulty. If a fault occurs, recursive analysis continues until the root cause of the fault is determined. If no fault occurs, the target service is directly identified as the root cause of the fault. This adds a layer of protection mechanism to the basic positioning process, which can effectively avoid misjudgment of faults caused by incomplete information, complex dependencies, or one-sided analysis.

[0129] Figure 3 A flowchart illustrating another embodiment of the fault location method provided in this application is shown below. Figure 1 Based on the illustrated process, this section mainly describes how to generate a fault analysis report. (See [link to relevant documentation]). Figure 3 As shown, it includes the following steps:

[0130] Step 301: Receive fault alarm information from the microservice system.

[0131] Step 302: Using the pre-trained first language model, based on the topology of the microservice system, perform semantic analysis on the fault alarm information to determine the target service associated with the fault.

[0132] Step 303: Obtain the operation-related data of the target service.

[0133] Step 304: Use the pre-trained second language model to perform reasoning analysis on the operation-related data of the target service to obtain the cause of the target service failure.

[0134] Step 305: Determine whether the failure of the target service is caused by a failure of the upstream service of the target service. If it is determined that the failure of the target service is caused by a failure of the upstream service of the target service, proceed to step 306; if it is determined that there is no upstream service causing the failure, proceed to step 307.

[0135] Step 306: Update the upstream service to the new target service and return to step 303.

[0136] Step 307: Obtain the root cause of the fault.

[0137] For the relevant descriptions of steps 301-307 above, please refer to the above. Figure 1 Detailed description of the relevant embodiments.

[0138] Step 308: After obtaining the root cause of the fault, generate a fault analysis report based on the root cause.

[0139] In one embodiment, after obtaining the root cause of the fault, the specific implementation of generating a fault analysis report based on the root cause is as follows: based on the root cause of the fault and the fault cause of each target service determined in the entire fault location process, a fault propagation path is generated, and a pre-trained third language model is used to perform reasoning analysis on the root cause of the fault and the fault propagation path to obtain a repair suggestion for the root cause of the fault; the fault propagation path, the root cause of the fault, and the repair suggestion for the root cause of the fault are integrated into a fault analysis report in a preset format.

[0140] The above-mentioned remediation recommendations include at least one of the following: service restart, request rate limiting, resource expansion, or code optimization.

[0141] The preset format can be a pre-defined report output format, such as PDF format, Word text, or chart format. This application embodiment does not limit this.

[0142] For example, based on the fault causes and timelines of each target service in steps 301-307, the fault propagation path is traced according to the root cause-downstream service logic. The fault service, fault manifestation, and trigger time of each link are identified. The root cause and fault propagation path are input into the pre-trained third language model. Three types of suggestions are proposed for the root cause: emergency repair, short-term optimization, and long-term prevention, yielding the model's output. Following a pre-set enterprise fault analysis report template (including title, fault overview, propagation path, root cause, and root cause analysis), the above information is integrated into a structured report output.

[0143] pass Figure 3 The description of the illustrated embodiment systematically outlines fault propagation paths and fault repair suggestions, integrating them into standardized reports. This achieves full-process coverage of fault location, from technical investigation to root cause analysis and repair recommendations, enhancing the team's fault response capabilities and enabling rapid identification and repair of fault root causes. Furthermore, it enables the visualization of fault information, supports fault review, and allows for targeted optimization of the system based on the review results, thereby improving system stability and reliability.

[0144] Figure 4 A block diagram illustrating an embodiment of a fault location device provided in this application, as shown below. Figure 4 As shown, the device includes:

[0145] Alarm information receiving module 41 is used to receive fault alarm information sent by the microservice system;

[0146] The target service determination module 42 is used to perform semantic analysis on the fault alarm information based on the topology of the microservice system using a pre-trained first language model to determine the target service associated with the fault.

[0147] Data acquisition module 43 is used to acquire operation-related data of the target service;

[0148] The fault cause acquisition module 44 is used to perform reasoning analysis on the operation-related data of the target service using a pre-trained second language model to obtain the fault cause of the target service.

[0149] The fault root cause determination module 45 is used to update the upstream service to a new target service when it is determined that the fault of the target service is caused by the fault of the upstream service of the target service, and return to execute the steps of obtaining the operation-related data of the target service and thereafter, until it is determined that there is no upstream service causing the fault, and thus obtain the fault root cause.

[0150] In one possible implementation, the target service determination module 42 is specifically used for:

[0151] The fault alarm information and the topology of the microservice system are input into a pre-trained first language model, so that the first language model can perform semantic analysis on the fault alarm information based on the topology of the microservice system to determine the target service associated with the fault.

[0152] In one possible implementation, the data acquisition module 43 is specifically used for:

[0153] Call at least one of the distributed tracing interface, service performance metric interface, and log data interface of the target service to obtain the operation-related data of the target service; wherein, the operation-related data includes at least one of the following: distributed tracing data, performance monitoring data, and log monitoring data.

[0154] In one possible implementation, determining that the failure of the target service was caused by a failure of an upstream service of the target service includes:

[0155] Based on the topology of the microservice system, a set of upstream services that have a direct calling relationship with the target service is determined;

[0156] The abnormal features in the operation-related data of each upstream service in the upstream service set are extracted by the pre-trained third language model, and the abnormal features are matched with the failure cause of the target service by time sequence and causal correlation degree calculation.

[0157] If the anomaly of any upstream service occurs earlier than the time of the fault cause of the target service, and the causal correlation exceeds a preset threshold, then the fault cause of the target service is determined to be caused by the fault of that upstream service.

[0158] In one possible implementation, the device is further used for:

[0159] If it is determined that there is no upstream service causing the failure for the initially identified target service, candidate upstream services that meet the set conditions are selected from the set of first-level upstream services that have a direct calling relationship with the initially identified target service.

[0160] Obtain the operation-related data of the candidate upstream services;

[0161] The second language model is used to perform reasoning analysis on the operation-related data of the candidate upstream service to determine whether the candidate upstream service is faulty.

[0162] If it is determined that the candidate upstream service is faulty, the candidate upstream service is updated to a new target service, and the process of obtaining the operation-related data of the target service and subsequent steps is returned until it is determined that there is no upstream service causing the fault, and the root cause of the fault is obtained.

[0163] If it is determined that the candidate upstream service is not faulty, the fault cause of the target service initially identified is determined as the root cause of the fault.

[0164] In one possible implementation, the device further includes:

[0165] The fault analysis report generation module is used to generate a fault analysis report based on the root cause of the fault after obtaining the root cause of the fault.

[0166] In one possible implementation, the fault analysis report generation module is specifically used for:

[0167] Based on the root cause of the fault and the fault cause of each target service determined in the entire fault localization process, a fault propagation path is generated, and a pre-trained third language model is used to perform reasoning analysis on the root cause of the fault and the fault propagation path to obtain a repair suggestion for the root cause of the fault.

[0168] The fault propagation path, the root cause of the fault, and the repair suggestions for the root cause are integrated into a fault analysis report in a preset format.

[0169] like Figure 5 As shown in the figure, this application provides an electronic device, including a processor 111, a communication interface 112, a memory 113, and a communication bus 114, wherein the processor 111, the communication interface 112, and the memory 113 communicate with each other through the communication bus 114.

[0170] Memory 113 is used to store computer programs;

[0171] In one embodiment of this application, when the processor 111 executes the program stored in the memory 113, it implements the fault location method provided in any of the foregoing method embodiments, including:

[0172] Receive fault alarm information from the microservice system;

[0173] Using a pre-trained first language model, based on the topology of the microservice system, semantic analysis is performed on the fault alarm information to determine the target service associated with the fault.

[0174] Obtain the operation-related data of the target service;

[0175] The second pre-trained language model is used to perform inference analysis on the operation-related data of the target service to obtain the cause of the target service's failure.

[0176] If it is determined that the failure of the target service is caused by a failure of the upstream service of the target service, the upstream service is updated to a new target service, and the process of obtaining the operation-related data of the target service and subsequent steps is repeated until it is determined that there is no upstream service causing the failure, thus obtaining the root cause of the failure.

[0177] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the fault location method provided in any of the foregoing method embodiments.

[0178] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0179] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software plus a general-purpose hardware platform, or of course, using hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0180] It should be understood that the terminology used herein is for the purpose of describing particular exemplary embodiments only and is not intended to be limiting. Unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “described” as used herein may also mean including the plural forms. The terms “comprising,” “including,” “containing,” and “having” are inclusive and therefore indicate the presence of the stated features, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not construed as requiring them to be performed in a particular order described or illustrated unless the order of performance is explicitly indicated. It should also be understood that additional or alternative steps may be used.

[0181] The above description is merely a specific embodiment of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A fault location method characterized by, The method comprises: receiving a fault alarm information issued by a microservice system; performing semantic analysis on the fault alarm information based on the topology structure of the microservice system by using a pre-trained first large language model to determine a target service associated with the fault; obtaining running related data of the target service; performing inference analysis on the running related data of the target service by using a pre-trained second large language model to obtain a fault cause of the target service; in a case where it is determined that the fault cause of the target service is caused by a fault of an upstream service of the target service, updating the upstream service to a new target service and returning to perform the step of obtaining the running related data of the target service and the subsequent steps until it is determined that there is no upstream service causing the fault, to obtain a root cause of the fault.

2. The method of claim 1, wherein, The method further comprises: in a case where it is determined that there is no upstream service causing the fault for the first determined target service, selecting a candidate upstream service satisfying a set condition from a first-level upstream service set having a direct calling relationship with the first determined target service; 3. The method of claim 1, wherein, obtaining running related data of the candidate upstream service; performing inference analysis on the running related data of the candidate upstream service by using the second large language model to determine whether the candidate upstream service has a fault; 4. The method of claim 1, wherein, in a case where it is determined that the candidate upstream service has a fault, updating the candidate upstream service to a new target service and returning to perform the step of obtaining the running related data of the target service and the subsequent steps until it is determined that there is no upstream service causing the fault, to obtain a root cause of the fault. The method further comprises: in a case where it is determined that there is no upstream service causing the fault for the first determined target service, selecting a candidate upstream service satisfying a set condition from a first-level upstream service set having a direct calling relationship with the first determined target service; obtaining running related data of the candidate upstream service; 5. The method according to any of claims 1 to 4, characterized in that, performing inference analysis on the running related data of the candidate upstream service by using the second large language model to determine whether the candidate upstream service has a fault; in a case where it is determined that the candidate upstream service has a fault, updating the candidate upstream service to a new target service and returning to perform the step of obtaining the running related data of the target service and the subsequent steps until it is determined that there is no upstream service causing the fault, to obtain a root cause of the fault. ​ ​ In a case where it is determined that the candidate upstream service exists a fault, the candidate upstream service is updated as a new target service, and the steps of obtaining running related data of the target service and the following steps are performed until it is determined that there is no upstream service causing a fault, and a fault root cause of the fault is obtained; In a case where it is determined that the candidate upstream service does not exist a fault, the fault cause of the target service determined for the first time is determined as the fault root cause of the fault.

6. The method of claim 1, wherein, The method further comprises: After the fault root cause of the fault is obtained, a fault analysis report is generated according to the fault root cause.

7. The method of claim 6, wherein, The generating of the fault analysis report according to the fault root cause comprises: According to the fault root cause and the fault cause of each target service determined in the entire fault locating process, a fault propagation path is generated, and the fault root cause and the fault propagation path are analyzed by reasoning using a pre-trained third large language model, to obtain a repair suggestion of the fault root cause; The fault propagation path, the fault root cause and the repair suggestion of the fault root cause are integrated into a fault analysis report in a preset format.

8. A fault location device characterized by, The apparatus comprises: An alarm information receiving module configured to receive fault alarm information sent by a microservice system; A target service determining module configured to perform semantic analysis on the fault alarm information based on a topology structure of the microservice system using a pre-trained first large language model, to determine a target service associated with a fault; A data obtaining module configured to obtain running related data of the target service; A fault cause obtaining module configured to perform reasoning analysis on the running related data of the target service using a pre-trained second large language model, to obtain a fault cause of the target service; A fault root cause determining module configured to, in a case where it is determined that the fault cause of the target service is caused by a fault of an upstream service of the target service, update the upstream service as a new target service, and return to perform the steps of obtaining the running related data of the target service and the following steps until it is determined that there is no upstream service causing a fault, and a fault root cause of the fault is obtained.

9. An electronic device, comprising: Comprise: A processor and a memory, the processor is used for executing a fault locating control program stored in the memory, to realize the fault locating method in any one of claims 1-7.

10. A storage medium, characterized by The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to realize the fault locating method in any one of claims 1-7.