Fault root cause positioning method and device and storage medium

By constructing a service node call relationship link graph and using graph algorithms to analyze single-step request timeout data across service nodes, the problem of traditional fault monitoring being unable to accurately locate the root cause of distributed system failures is solved, and the accurate identification and automatic location of the service node that is the root cause of the failure is achieved.

CN120614243APending Publication Date: 2025-09-09CHINA MERCHANTS BANK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510761816.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

In a distributed system architecture, traditional fault monitoring methods cannot accurately locate the root cause of faults across service call links, making it difficult to identify the true source of the fault, affecting system stability and availability.

Method used

By receiving the thread pool full signal sent by the message middleware, we enter the fault root cause location phase, build a service node call relationship link graph, use the graph algorithm to analyze the single-step request timeout data across service nodes, and determine the service node that is the root cause of the fault.

Benefits of technology

It achieves accurate identification of the root causes of distributed system resource contention failures, improves the accuracy and automation level of troubleshooting, and overcomes the root cause tracing limitations caused by data fragmentation in traditional solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120614243A_ABST
    Figure CN120614243A_ABST
Patent Text Reader

Abstract

The invention discloses a fault root cause positioning method and device and a storage medium, and relates to the technical field of fault processing, and the method comprises the steps: receiving a thread pool full signal sent by message-oriented middleware, and entering a fault root cause positioning stage; constructing a service node calling relation link diagram according to at least one cross-service node single-step request timeout data generated in the fault root cause positioning stage; according to the service node calling relation link diagram, the fault root cause service node is determined, the problem that in the prior art, it is difficult to accurately locate the distributed system resource competition fault root cause service node is solved, and accurate recognition of the fault root cause service node is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of fault handling technology, and in particular to a method, device, and storage medium for locating the root cause of a fault. Background Art

[0002] In a distributed system architecture, the call relationships between service nodes form a mesh topology. A single node anomaly can trigger cascading failures across services, making troubleshooting increasingly difficult. In a complex microservice architecture, limited resources can lead to cascading timeouts and resource contention, resulting in failures. Traditional fault monitoring methods can only detect local anomalies and lack the ability to globally analyze cross-service call chains. Related technologies rely on manual experience or fragmented log information to perform root cause identification, making it impossible to accurately locate the service node at the root of distributed system resource contention failures. This makes it difficult to accurately identify the true source of the failure, affecting system stability and availability. Summary of the Invention

[0003] The main purpose of this application is to provide a method, device and storage medium for locating the root cause of a fault, aiming to solve the technical problem in the prior art that the root service node of a distributed system resource competition fault cannot be accurately located.

[0004] To achieve the above objectives, the present application proposes a fault root cause location method, which is applied to a fault root cause location system. The fault root cause location method includes:

[0005] Receive the thread pool full signal sent by the message middleware and enter the fault root cause identification phase;

[0006] Constructing a service node call relationship link graph based on at least one cross-service node single-step request timeout data generated in the fault root cause locating phase;

[0007] The service node that is the root cause of the fault is determined according to the service node call relationship link diagram.

[0008] In some embodiments, before the step of constructing a service node call relationship link graph based on at least one cross-service node single-step request timeout data generated in the fault root cause locating stage, the step further includes:

[0009] Receive the connection pool full signal sent by the message middleware and enter the fault root cause location phase.

[0010] In some embodiments, after entering the fault root cause location phase, the method further includes:

[0011] At the end of each cycle in the fault root cause location phase, executing the steps of constructing a service node call relationship link graph based on at least one cross-service node single-step request timeout data generated in the fault root cause location phase; and determining the fault root cause service node based on the service node call relationship link graph;

[0012] If the thread pool full signal is not detected within a preset time, the root cause location phase ends.

[0013] In some embodiments, the step of constructing a service node call relationship link graph based on at least one cross-service node request timeout data generated during the fault root cause location phase includes:

[0014] Extracting the request identification information, single-step request identification information, requesting service node identification information, requested service node identification information, single-step request response time, single-step request response timeout time, and single-step request result from each cross-service node single-step request timeout data;

[0015] Aggregating cross-service node single-step request timeout data belonging to the same request based on each of the request identification information;

[0016] For each request, reconstruct the call path between each service node based on each single-step request identification information, each requesting service node identification information, and each requested service node identification information to generate an initial call tree with a time sequence relationship;

[0017] Obtaining a service node initial call relationship link graph based on at least one generated initial call tree, wherein the service node initial call relationship link graph includes a plurality of vertices and a plurality of edges connecting the vertices, wherein the vertices are used to represent the service nodes, and the edges are used to represent the call relationship between two service nodes connected by the edges;

[0018] Determine the timeout propagation weight of each edge based on each single-step request response time and each single-step request response timeout;

[0019] The service node call relationship link graph is determined according to the timeout propagation weight of each edge and the service node initial call relationship link graph.

[0020] In some embodiments, the step of determining the timeout propagation weight of each edge based on each single-step request response time and each single-step request response timeout time includes:

[0021] For each timeout event, determining a timeout ratio corresponding to the edge in the timeout event according to the single-step request response time and the single-step request response timeout time;

[0022] At least one of the timeout ratios of the edge is summed to obtain a timeout propagation weight of the edge.

[0023] In some embodiments, the step of determining the service node that is the root cause of the fault based on the service node call relationship link diagram includes:

[0024] Using a graph algorithm to perform centrality analysis on each of the service nodes in the service node call relationship link graph to obtain an importance score for each service node;

[0025] Based on the importance scores of the service nodes, the service node corresponding to the service node with the largest importance score is determined as the fault root cause service node.

[0026] In some embodiments, after the step of determining the service node that is the root cause of the fault based on the service node call relationship link diagram, the following steps are included:

[0027] Obtaining log data of the fault root cause service node, and determining a fault category of the fault root cause service node according to the log data;

[0028] Determining a target repair strategy corresponding to the fault category based on a correspondence between the fault category and pre-set candidate fault categories and fault repair strategies;

[0029] The target repair strategy is invoked to repair the fault of the root cause service node.

[0030] In addition, to achieve the above objectives, the present application also proposes a method for locating the root cause of a fault, which is applied to a distributed system. The method for locating the root cause of a fault includes:

[0031] Monitoring the resource status of the worker thread pool in each service node, and generating a thread pool full signal if the worker thread pool is determined to be full based on the resource status of the worker thread pool;

[0032] Monitoring the single-step request processing results of each of the service nodes;

[0033] When it is detected that the single-step request response time in the single-step request processing result is greater than the corresponding time threshold, generating cross-service node single-step request timeout data according to the single-step request processing result;

[0034] The thread pool full signal and the cross-service node single-step request timeout data are sent to the message middleware.

[0035] In some embodiments, the root cause location method for the fault applied to the distributed system further includes:

[0036] Monitor the resource status of the database connection pool;

[0037] If it is determined based on the resource status of the database connection pool that the database connection pool is in the pool full state, generating a connection pool full signal;

[0038] The connection pool full signal is sent to the message middleware.

[0039] In addition, to achieve the above-mentioned purpose, the present application also proposes a device for locating the root cause of a fault, the device comprising:

[0040] The pool full signal identification module is used to receive the thread pool full signal sent by the message middleware and enter the fault root cause location phase;

[0041] A link graph construction module, configured to construct a service node call relationship link graph based on at least one cross-service node single-step request timeout data generated during the fault root cause location phase;

[0042] The root cause location module is used to determine the service node that is the root cause of the fault based on the service node call relationship link diagram.

[0043] In addition, to achieve the above-mentioned purpose, the present application also proposes a device for locating the root cause of a fault, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the method for locating the root cause of a fault as described above.

[0044] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium and stores a computer program. When the computer program is executed by a processor, it implements the steps of the root cause locating method of the fault as described above.

[0045] In addition, to achieve the above-mentioned object, the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps of the method for locating the root cause of a fault as described above.

[0046] One or more technical solutions proposed in this application have at least the following technical effects: by receiving the thread pool full signal sent by the message middleware, when the thread pool full signal is detected, the fault root cause location phase is entered, and the fault analysis is automatically triggered based on the thread pool full signal. In the fault root cause location phase, single-step request timeout data across service nodes is collected, and a service node call relationship link graph including service dependencies, timing characteristics, and resource competition status is constructed based on at least one single-step request timeout data across service nodes generated in the fault root cause location phase. Based on at least one single-step request timeout data across service nodes generated in the fault root cause location phase, a service node call relationship link graph is constructed. Through the service node call relationship link graph, the dependency relationships and request flow paths between services can be clearly visualized, providing a basis for subsequent root cause analysis of the fault. An in-depth analysis is conducted based on the service node call relationship link graph to determine the service node that is the root cause of the fault. This application uses the single-step request timeout data across service nodes in the fault root cause location stage to construct a global service node call relationship chain graph, overcoming the root cause tracing limitations of traditional solutions due to data fragmentation, so that fault analysis is not limited to a single node, but covers all related services in the entire distributed system. Analysis based on the service node call relationship chain graph can more accurately find the real root cause of the fault and obtain the root cause service node of the fault, solving the problem in the existing technology that it is difficult to accurately locate the root cause service node of the distributed system resource competition fault, realizing accurate identification of the root cause service node of the fault, and improving the accuracy and automation level of fault troubleshooting. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0048] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0049] Figure 1 A flowchart illustrating a method for locating the root cause of a fault according to the present invention;

[0050] Figure 2 A flowchart illustrating a second embodiment of the method for locating the root cause of a fault in this application is provided;

[0051] Figure 3 Schematic diagram of resource contention scenario provided for the root cause location method of the fault in this application;

[0052] Figure 4This is a schematic diagram of the module structure of the device for locating the root cause of a fault according to an embodiment of the present application;

[0053] Figure 5 Schematic diagram of the device structure of the hardware operating environment involved in the root cause locating method of the fault in the embodiment of the present application. DETAILED DESCRIPTION

[0054] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.

[0055] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0056] The main solution of the embodiment of the present application is: receiving the thread pool full signal sent by the message middleware and entering the fault root cause location stage; constructing a service node call relationship link graph based on at least one cross-service node single-step request timeout data generated in the fault root cause location stage; determining the fault root cause service node based on the service node call relationship link graph.

[0057] In this embodiment, for ease of description, the following description is made with the fault monitoring system as the execution subject.

[0058] Because existing technologies in distributed system architectures employ a mesh-like topology for call relationships between service nodes, a single node failure can trigger cascading failures across services, making troubleshooting increasingly difficult. In complex microservice architectures, limited resources can lead to cascading timeouts and resource contention, resulting in failures. Traditional fault monitoring methods can only detect local anomalies and lack the ability to globally analyze cross-service call chains. Existing root cause identification relies on manual experience or fragmented log information, making it impossible to accurately locate the service node responsible for distributed system resource contention failures. This makes it difficult to accurately identify the true source of the failure, impacting system stability and availability.

[0059] The present application provides a solution. The present application uses the single-step request timeout data across service nodes in the fault root cause location stage to construct a global service node call relationship link graph, overcoming the root cause tracing limitations of traditional solutions caused by data fragmentation, so that fault analysis is not limited to a single node, but covers all related services in the entire distributed system. Analysis based on the service node call relationship link graph can more accurately find the real root cause of the fault and obtain the root cause service node of the fault, solving the problem in the existing technology that it is difficult to accurately locate the root cause service node of the distributed system resource competition fault, realizing accurate identification of the root cause service node of the fault, and improving the accuracy and automation level of fault troubleshooting.

[0060] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, mobile phone, etc., or a fault root cause location device capable of performing the above functions. The following describes this embodiment and the following embodiments using a fault monitoring system as an example.

[0061] Based on this, the embodiment of the present application provides a method for locating the root cause of a fault, referring to Figure 1 and Figure 2 , Figure 1 and Figure 2 This is a flow chart of the first embodiment of the fault root cause location method of the present application. The first embodiment is executed by a monitoring system, which includes a distributed system, a fault root cause location system and a message middleware. Figure 1 The fault root cause location method process is executed by the fault root cause location system; Figure 2 The root cause location method process of the fault is executed by the distributed system.

[0062] In the first embodiment, the method for locating the root cause of a fault includes steps 101 to 103 and steps 201 to 203 .

[0063] Step 101: Receive a thread pool full signal sent by the message middleware and enter the fault root cause location phase.

[0064] Specifically, a thread pool full signal triggers the root cause location phase, indicating that a service node's worker thread pool has reached its resource usage limit, potentially leading to cascading failures. The root cause location phase automatically begins fault analysis after detecting a distributed system anomaly (such as a thread pool full signal or a connection pool full signal). This phase collects, integrates, and analyzes single-step request timeout data across service nodes to pinpoint the service node responsible for the failure.

[0065] In some embodiments, a thread pool full signal pushed by the message middleware is received. When the thread pool full signal is detected, the fault root cause location phase is entered. During the fault root cause location phase, a service node call relationship link diagram is constructed based on at least one cross-service node single-step request timeout data generated during the fault root cause location phase. In addition, log data and performance indicators of each service node are collected for comprehensive analysis. The performance indicators of the service node may include central processing unit (CPU) utilization, memory usage, response time, etc.

[0066] Step 102: construct a service node call relationship link graph based on at least one cross-service node single-step request timeout data generated in the fault root cause location phase.

[0067] Specifically, in a distributed system, when a service node initiates a request to another service node, if the request processing time exceeds a preset time threshold, relevant information about the request is recorded. The relevant information may include but is not limited to request identification information, single-step request identification information, requesting service node identification information, requested service node identification information, single-step request response time, single-step request response timeout time, single-step request result, etc., and the relevant information of the request is encapsulated as cross-service node single-step request timeout data. Cross-service node request timeout data can be obtained by applying the following steps to a distributed system: monitoring the single-step request processing results of each service node; when it is detected that the single-step request response time in the single-step request processing result is greater than the corresponding time threshold, generating cross-service node single-step request timeout data based on the single-step request processing result.

[0068] As an example, at least one cross-service node single-step request timeout data generated during the fault root cause location phase is obtained from the message middleware. The cross-service node single-step request timeout data includes request identification information, single-step request identification information, requesting service node identification information, requested service node identification information, single-step request response time, single-step request response timeout time, single-step request result, etc. The cross-service node single-step request timeout data is cleaned and preprocessed to remove duplicate data and handle missing values. A graph database or graph computing framework can be used to construct a weighted service node call relationship link graph with service nodes as vertices and call relationships as edges. The service node call relationship link graph can intuitively display the propagation path of each request between service nodes. The root node that causes cascade blockage can be located by reversely tracing high-weight edges, solving the defect that traditional single-point analysis cannot associate cross-service dependencies, and helping to accurately identify the source of resource competition.

[0069] Step 103: Determine the service node that is the root cause of the fault based on the service node call relationship link diagram.

[0070] In some embodiments, a graph algorithm can be used to identify nodes with high centrality. Nodes with high centrality are nodes that occupy key positions in the network structure, and these nodes are likely to be the core points of fault propagation. The graph algorithm can be a shortest path algorithm, PageRank, etc. In combination with the real-time log data of each service node (such as CPU usage, memory usage, request success rate, etc.), it is further determined whether the above-mentioned high-centrality nodes do have resource competition or performance bottleneck problems. In addition, the data differences during normal operation and when the fault occurs can be compared to determine which nodes have significantly changed their behavior. By using the service node call relationship link diagram, the root cause service node that caused the fault can be determined to improve the accuracy of fault location.

[0071] Based on the fault root cause location method provided by the present application, by receiving the thread pool full signal sent by the message middleware, when the thread pool full signal is detected, the fault root cause location phase is entered, and the fault analysis is automatically triggered based on the thread pool full signal. In the fault root cause location phase, single-step request timeout data across service nodes is collected, and a service node call relationship link graph including service dependencies, timing characteristics, and resource competition status is constructed based on at least one single-step request timeout data across service nodes generated in the fault root cause location phase. Based on at least one single-step request timeout data across service nodes generated in the fault root cause location phase, a service node call relationship link graph is constructed. Through the service node call relationship link graph, the dependency relationships and request flow paths between services can be clearly visualized, providing a basis for subsequent root cause analysis of the fault. An in-depth analysis is conducted based on the service node call relationship link graph to determine the service node that is the root cause of the fault. This application uses the single-step request timeout data across service nodes in the fault root cause location stage to construct a global service node call relationship chain graph, overcoming the root cause tracing limitations of traditional solutions due to data fragmentation, so that fault analysis is not limited to a single node, but covers all related services in the entire distributed system. Analysis based on the service node call relationship chain graph can more accurately find the real root cause of the fault and obtain the root cause service node of the fault, solving the problem in the existing technology that it is difficult to accurately locate the root cause service node of the distributed system resource competition fault, realizing accurate identification of the root cause service node of the fault, and improving the accuracy and automation level of fault troubleshooting.

[0072] Step 201 : monitor the resource status of the work thread pool in each service node. If the work thread pool is determined to be full based on the resource status of the work thread pool, generate a thread pool full signal.

[0073] Service nodes are independently running service units or components in a distributed system. Each service node may be responsible for performing specific functions and interacting with other service nodes. Each service node has independent computing resources and software resources. Computing resources can include CPUs and memory, while software resources can include thread pools and connection pools. Service nodes can be application software. A worker thread pool is a resource pool used to manage multiple worker threads within a service node. It is responsible for creating, destroying, and reusing threads, improving the system's efficiency in handling concurrent requests. By creating a worker thread pool, the overhead of frequent thread creation and destruction can be avoided, resulting in more efficient use of system resources. The resource status of a worker thread pool reflects the thread resource usage within the pool, including metrics such as the number of active threads (the number of threads currently executing tasks), the queue size (the length of the queue waiting for tasks), and the maximum number of threads (the maximum number of threads the pool can accommodate). When the number of active threads in a worker thread pool is greater than or equal to the maximum number of threads, and the queue is full and cannot accept new tasks, the pool is considered full. In this state, newly submitted tasks cannot be immediately executed, potentially resulting in a denial of service. The thread pool full signal is a notification or alarm signal generated when the worker thread pool is detected to be full. The thread pool full signal is used to trigger the fault root cause location process and enter the fault root cause location phase to promptly identify and resolve the problem that caused the thread pool to be full.

[0074] In some embodiments, key performance indicators of the work thread pool in each service node can be regularly collected through monitoring tools (such as Prometheus, Micrometer, etc.). The key performance indicators of the work thread pool may include the number of active threads, the maximum number of threads, the length of the waiting queue, etc., and the key performance indicators of the work thread pool are used as the resource status of the work thread pool. According to the resource status of the work thread pool, a preset threshold or algorithm can be used to judge the working status of the work thread pool. For example, when the number of active threads continues to reach the maximum number of threads and the queue size exceeds a certain threshold, the state of the work thread pool can be determined to be a pool full state. When it is determined that the thread pool is in a pool full state, a standardized thread pool full signal is generated, which can be published to the message middleware through Kafka or RabbitMQ.

[0075] Step 202 : monitor the single-step request processing results of each service node. When it is detected that the single-step request response time in the single-step request processing result is greater than the corresponding time threshold, generate cross-service node single-step request timeout data according to the single-step request processing result.

[0076] Specifically, the single-step request processing result refers to the detailed information about the request processed by each independent service node in a distributed system, including request identification information, source service node identification information, target service node identification information, requesting service node identification information, requested service node identification information, request start time, request end time, single-step request response time, single-step request response timeout period, and single-step request result (success / failure / timeout). The single-step request response time refers to the actual time it takes from a service node to send a request and receive a response. The time threshold corresponding to the single-step request processing time is a pre-set maximum allowable response time for each request type. If the actual single-step request response time of a request exceeds the corresponding threshold, the request is considered to have timed out. Cross-service node single-step request timeout data is structured data that records request processing timeouts caused by resource competition at a service node in a distributed call chain. Single-step request timeout data includes request identification information, single-step request identification information, requesting service node identification information, requested service node identification information, single-step request response time, single-step request response timeout period, and single-step request result.

[0077] In some embodiments, distributed tracing tools can be deployed to capture the single-step request processing results of each service node in real time, including key data such as request identification information, source service node identification information, target service node identification information, request service node identification information, requested service node identification information, request start time, request end time, single-step request response time, single-step request response timeout time, single-step request result (success / failure / timeout), etc. When it is detected that the actual processing time of a single-step request exceeds a preset time threshold (for example, a query request is considered to have timed out if it exceeds 200 milliseconds), the request is automatically marked as a timed-out request, and cross-service node single-step request timeout data is generated based on the single-step request processing results, providing a data basis for the subsequent construction of a service node call relationship link diagram and fault root cause location.

[0078] Step 203: Send the thread pool full signal and cross-service node single-step request timeout data to the message middleware.

[0079] In some embodiments, in a distributed system, in order to improve the observability and fault response capability of the system, when a thread pool full signal and / or a cross-service node single-step request timeout is generated, the thread pool full signal and the cross-service node single-step request timeout data are sent to the message middleware. The API provided by the message middleware can be used to asynchronously send the thread pool full signal and the cross-service node single-step request timeout data to a specific topic or queue. For example, an efficient message delivery mechanism is implemented using tools such as Kafka or RabbitMQ to ensure reliable message transmission even under high load conditions. The fault root cause location system or other specialized services subscribe to and consume the messages for further processing, such as triggering alarms, recording logs, or automatically executing repair strategies.

[0080] Based on the first embodiment of the present application, in the second embodiment of the present application, the root cause location method of a fault applied to a distributed system further includes:

[0081] Monitor the resource status of the database connection pool;

[0082] If the database connection pool is determined to be full based on the resource status of the database connection pool, a connection pool full signal is generated;

[0083] Send a connection pool full signal to the message middleware.

[0084] Specifically, the database connection pool is a resource pool that manages database connections. It is responsible for creating, reusing, and destroying connections to avoid the overhead of frequently establishing connections. Building a database connection pool can improve database access performance. The resource status of the database connection pool is the occupancy of connection resources in the database connection pool, including the number of active connections (the number of connections in use), the number of idle connections (the number of connections that can be immediately allocated), the maximum number of connections (the upper limit of the connection pool capacity), and the waiting queue length (the number of requests waiting for available connections). The pool full state is the state of the database connection pool when the number of active connections reaches the maximum number of connections and the waiting queue is full. When the database connection pool is in the pool full state, new requests cannot obtain database connections. The connection pool full signal is an alarm signal generated when the database connection pool is detected to be in the pool full state. It is used to trigger the fault root cause location process, indicating that database resources may become a performance bottleneck.

[0085] As an example, key performance indicators of the database connection pool in each service node can be regularly collected through monitoring tools (such as Prometheus, Micrometer, etc.). The key performance indicators of the database connection pool may include the number of active connections, the maximum number of connections, and the length of the waiting queue, etc. The key performance indicators of the database connection pool are used as the resource status of the database connection pool. According to the resource status of the database connection pool, a preset threshold or algorithm can be used to determine the working status of the database connection pool. For example, when the number of active connections reaches the maximum number of connections and the length of the waiting queue exceeds a certain threshold, the status of the database connection pool can be determined to be a pool-full state. When it is determined that the thread pool is in a pool-full state, a standardized connection pool full signal is generated, which can be published to the message middleware through Kafka or RabbitMQ.

[0086] According to the first embodiment of the present application, before the step of constructing a service node call relationship link graph based on at least one cross-service node single-step request timeout data generated in the fault root cause locating phase, the step further includes:

[0087] Receive the connection pool full signal sent by the message middleware and enter the fault root cause identification phase.

[0088] For example, a connection pool full signal is generated when a database connection pool is full. This signal triggers the root cause identification process, indicating that database resources may be a performance bottleneck. The root cause identification phase automatically begins fault analysis after a system anomaly (connection pool full signal) is detected. This phase collects, integrates, and analyzes single-step request timeout data across service nodes to pinpoint the service node that caused the fault.

[0089] In some embodiments, a connection pool full signal pushed by the message middleware is received. Upon detecting the connection pool full signal, the fault root cause determination phase begins. During this phase, a service node call relationship diagram is constructed based on at least one cross-service node single-step request timeout data generated during the root cause determination phase. Furthermore, log data and performance metrics (such as CPU usage, memory usage, response time, etc.) are collected from each service node for comprehensive analysis.

[0090] In some embodiments, after receiving a thread pool full signal sent by the message middleware and entering the fault root cause location phase, the process further includes:

[0091] At the end of each cycle in the fault root cause location phase, executing the steps of constructing a service node call relationship link graph based on at least one cross-service node single-step request timeout data generated in the fault root cause location phase; and determining the service node that is the root cause of the fault based on the service node call relationship link graph;

[0092] If the thread pool full signal is not detected within the preset time, the root cause location phase ends.

[0093] Specifically, when a thread pool full signal is detected, the fault root cause location phase is immediately initiated, and the root cause location analysis of the fault is performed based on the cross-service node single-step request timeout data. Within each cycle (e.g., 10 seconds) of the fault root cause location phase, cross-service node single-step request timeout data is received, including request identification information, requesting service node identification information, requested service node identification information, single-step request response time, etc. At the end of each cycle, based on all cross-service node single-step request timeout data obtained since the start of the fault root cause location phase, a service node call relationship link graph is constructed or updated, and a graph algorithm is used to perform centrality analysis on each service node in the service node call relationship link graph, and the node with the highest importance score is determined as the fault root cause service node. The duration of each cycle in the fault root cause location phase is consistent, but this application does not impose a specific duration limit on each cycle. Specifically, each cycle can be 10 seconds, 15 seconds, etc. The thread pool full signal is continuously detected. If the thread pool full signal is not detected within the preset time, it is determined that the system has resumed stable operation, the fault root cause location task is completed, and the root cause location phase is exited. By periodically updating the link graph and the above-mentioned dynamic exit mechanism, real-time tracking of the fault propagation path and resource optimization are achieved, avoiding long-term occupation of analysis resources. The deep positioning logic is activated only during the duration of the fault, improving system stability and resource utilization.

[0094] As an example, see Figure 3The distributed system includes multiple service nodes, including at least service node A, service node B, service node C, service node D, service node E, service node X, and service node Y. The resource status of the working thread pool in each service node and the resource status of the database connection pool in each service node are monitored. Based on the resource status of the working thread pool D, it is determined that the working thread pool D is in a pool full state, a thread pool full signal is generated, and the thread pool full signal is sent to the message middleware. The fault root cause locating system receives the thread pool full signal sent by the message middleware and enters the fault root cause locating phase. In the first cycle of the fault root cause locating phase, the cross-service node single-step request timeout data generated by the service node C is received. At the end of the first cycle of the fault root cause locating phase, based on the cross-service node single-step request timeout data obtained in the first cycle of the fault root cause locating phase, a service node call relationship link graph is constructed in the first cycle of the fault root cause locating phase. The service node call relationship link graph constructed in the first cycle includes service node B and service node C, as well as the call relationship between service node B and service node C. During the second cycle of the fault root cause location phase, cross-service node single-step request timeout data generated by service nodes B, D, and Y is received, as well as thread pool full signals generated by service nodes A and D and a connection pool full signal generated by service node X. At the end of the second cycle of the fault root cause location phase, a service node call relationship link graph is constructed for the second cycle of the fault root cause location phase based on the cross-service node single-step request timeout data obtained during the second cycle and the cross-service node single-step request timeout data obtained during the first cycle. The service node call relationship link graph constructed during the second cycle includes service nodes A, B, C, D, X, and Y, as well as the call relationships between the service nodes. After the first root cause tracing, the fault root cause location system will continue to collect new abnormal signals (pool full signals) and cross-service node single-step request timeout data generated by each service node to conduct a second or even Nth root cause tracing. At the end of the Nth cycle in the root cause location phase, a service node call relationship chain graph is constructed based on the cross-service node single-step request timeout data obtained from the first to the Nth cycle. The service node call relationship chain graph constructed in the Nth cycle includes service node A, service node B, service node C, service node D, service node X, service node Y, and service node E, as well as the call relationships between the service nodes. Each root cause tracing operation updates the service node call relationship chain graph and further confirms or adjusts the determination of the service node as the root cause of the fault.During the fault root cause location phase, if no thread pool full signal or connection pool full signal is detected within the preset time, it indicates that the current fault problem may have been resolved or temporarily stabilized, and the root cause location phase can be terminated.

[0095] In some embodiments, the step of constructing a service node call relationship link graph based on at least one cross-service node request timeout data generated during the fault root cause location phase includes:

[0096] Extract the request identification information, single-step request identification information, requesting service node identification information, requested service node identification information, single-step request response time, single-step request response timeout time, and single-step request result from each cross-service node single-step request timeout data;

[0097] Aggregate single-step request timeout data across service nodes belonging to the same request based on each request identification information;

[0098] For each request, based on the identification information of each single-step request, the identification information of each requesting service node, and the identification information of each requested service node, the call path between each service node is reconstructed to generate an initial call tree with a time sequence relationship;

[0099] Obtaining a service node initial call relationship link graph based on the generated at least one initial call tree, the service node initial call relationship link graph including a plurality of vertices and a plurality of edges connecting the vertices, wherein the vertices are used to represent service nodes, and the edges are used to represent the call relationship between two service nodes connected by the edges;

[0100] Determine the timeout propagation weight of each edge based on each single-step request response time and each single-step request response timeout;

[0101] Determine the service node call relationship link graph based on the timeout propagation weight of each edge and the service node initial call relationship link graph.

[0102] Specifically, the request identification information is an identifier used to uniquely identify a complete request. The single-step request identification information is an identifier that uniquely identifies a specific service node processing step in the call chain. The requesting service node identification information is the unique identifier of the service node that initiates the single-step request. The requested service node identification information is the unique identifier of the service node that receives and processes the single-step request. The single-step request response time is the time consumed by the service node to process the single-step request. The single-step request result is the processing result status of the single-step request, such as success, failure, or timeout. The timeout propagation weight is an indicator used to quantify the degree of propagation of the timeout in the call chain, and is related to the single-step request response time, the single-step request response timeout, and the request result.

[0103] As an example, cross-service node single-step request timeout data is received from the message middleware, and data is extracted for each cross-service node single-step request timeout data to obtain the request identification information, single-step request identification information, requesting service node identification information, requested service node identification information, single-step request response time, single-step request response timeout time, and single-step request result in the cross-service node single-step request timeout data. Based on the request identification information (TraceID), the cross-service node single-step request timeout data is grouped and aggregated, and timeout events belonging to the same call chain are associated together. For example, Python's defaultdict data structure can be used, with TraceID as the key, to aggregate cross-service node single-step request timeout data with the same TraceID into the same list. For each aggregated timeout call chain, the call path between the service nodes is reconstructed based on the time sequence of the single-step request identification information (SpanID). The entry service node of the call chain can be used as the root node, and downstream service nodes can be added as child nodes in the order of the call, thereby generating an initial call tree with a time sequence relationship. All generated initial call trees are merged to construct a service node initial call relationship chain graph. The service node initial call relationship chain graph is a directed graph consisting of multiple vertices and multiple edges. For each edge in the service node initial call relationship chain graph, a timeout propagation weight is calculated based on the response time and response timeout of each single-step request. This weight quantifies the extent to which timeouts propagate through the call chain. Vertices represent service nodes, and edges represent call relationships between service nodes. The calculated timeout propagation weights are annotated on the edges of the service node initial call relationship chain graph to generate the final service node call relationship chain graph, which is a weighted directed graph. By analyzing cross-service node request timeout data, a service node call relationship chain graph with timeout propagation weights is constructed to accurately locate the root cause of failures in distributed systems.

[0104] In some embodiments, the step of determining the timeout propagation weight of each edge based on each single-step request response time and each single-step request response timeout time includes:

[0105] For each timeout event, determine the timeout ratio of the edge corresponding to the timeout event based on the single-step request response time and the single-step request response timeout time;

[0106] The timeout propagation weight of the edge is obtained by summing up at least one timeout ratio of the edge.

[0107] Specifically, in distributed systems, the call chains between service nodes are complex and multi-layered. To determine which call paths between service nodes are most likely to cause failures and affect overall system stability, it is necessary to quantitatively analyze the timeout propagation weight of each call edge (i.e., the call relationship between two service nodes). When the actual response time of a service call (single-step request) exceeds a preset response time threshold, the service call is considered a timeout event. The timeout ratio is the ratio of the single-step request response timeout involved in a timeout event to the preset response time threshold (i.e., the single-step request response time minus the single-step request response timeout). If the single-step request response time exceeds the preset response time threshold, the timeout ratio can be used to measure the severity of the single-step request (timeout event). For example, if the single-step request response time in a timeout event is 2500 milliseconds and the single-step request response timeout is 500 milliseconds, the calculated timeout ratio in the timeout event is 0.25. The timeout propagation weight is a comprehensive indicator calculated based on the timeout ratio of an edge in multiple timeout events. It is used to quantify the timeout risk and impact of the edge in the entire system.

[0108] As an example, all cross-service node single-step request timeout data collected during the root cause location phase is traversed. Based on the request identification information (TraceID) and call relationship information (requesting service node identification information and requested service node identification information), an initial call tree with a temporal relationship is constructed. At least one of the generated initial call trees is merged to generate a service node call relationship link graph. For each timeout event, the timeout ratio of the edge in the timeout event is determined based on the single-step request response time and the single-step request response timeout. For example, if the single-step request response time in a timeout event is 2500 milliseconds and the single-step request response timeout is 500 milliseconds, the calculated timeout ratio in the timeout event is 0.25. For each edge, the timeout propagation weight of the edge is calculated by summing the timeout ratios in all related timeout events. For example, if the timeout ratios of an edge in three different timeout events are 0.2, 0.3, and 0.4, respectively, the timeout propagation weight of the edge is 0.9. Through quantitative analysis, abstract service call behaviors are transformed into measurable risk indicators, generating richer and more accurate service node call relationship diagrams, improving the accuracy of fault root cause location, and providing data support for automated diagnosis, resource scheduling optimization, and system robustness improvement.

[0109] In addition, when the actual timeout ratio of the edge in the timeout event calculated based on the single-step request response time and the single-step request response timeout time is greater than one, the value one is directly used as the timeout ratio of the edge in this timeout event.

[0110] In some embodiments, the step of determining the service node that is the root cause of the fault based on the service node call relationship link diagram includes:

[0111] Use graph algorithms to perform centrality analysis on each service node in the service node call relationship link graph to obtain the importance score of each service node;

[0112] Based on the importance scores of the service nodes, the service node corresponding to the one with the largest importance score is determined as the fault root cause service node.

[0113] As an example, based on a constructed service node call relationship graph, weighted directed graph data is extracted, including a set of service nodes, a set of call relationship edges, and the timeout propagation weight of each edge. A graph algorithm (such as PageRank, betweenness centrality, or eigenvector centrality) is used to calculate the importance of each service node. For example, when using the PageRank algorithm, the importance score of each service node can be iteratively calculated. Service nodes are sorted in descending order based on the importance scores output by the algorithm. A higher importance score indicates a stronger hub role in fault propagation. The service node with the highest importance score is selected as the root cause service node. If multiple service nodes have the same importance score, a secondary determination can be made by combining service topology or time series analysis. Using the constructed service node call relationship graph, a graph algorithm is used to perform centrality analysis on each service node in the graph to assess its importance within the entire call network. This allows the most likely root cause service node to be identified, accurately locating the hub node that triggers the cascading fault.

[0114] In some embodiments, after determining the service node that is the root cause of the fault based on the service node call relationship link diagram, the following steps are included:

[0115] Obtain log data of the fault root cause service node, and determine the fault type of the fault root cause service node based on the log data;

[0116] Determine the target repair strategy corresponding to the fault category based on the correspondence between the fault category and the pre-set candidate fault categories and fault repair strategies;

[0117] Invoke the target repair strategy to repair the fault of the service node that is the root cause of the fault.

[0118] Specifically, log data is recorded information generated by service nodes during operation. Log data can include system logs, application logs, and access logs. System logs include the entire operational status of the service node, including startup, shutdown, hardware errors, and kernel messages. Application logs include the service node's business logic processing and error messages. Access logs include access requests to the service node, including the request source, time, path, and response status code. Service node log data is automatically generated during operation and serves as an important basis for system monitoring, troubleshooting, performance optimization, and security audits. Fault categories are failure modes categorized based on log data analysis, such as database connection failures, thread pool exhaustion, configuration errors, code defects, and memory leaks. The candidate fault category-remediation strategy mapping is a predefined mapping between fault types and automated remediation actions. This mapping is developed by an experienced operations team based on historical data and best practices, covering various fault categories and their solutions. The target remediation strategy is an automated remediation solution matched from a mapping table for the current fault category, consisting of executable instructions or API call parameters.

[0119] As an example, after identifying the root cause service node, log data from the root cause service node can be retrieved from the log management system over the recent period. This log data contains information such as error messages, stack traces, and performance metrics, which can aid in further fault diagnosis. Furthermore, the fault category of the root cause service node can be identified based on the log data using standard text analysis techniques, pattern matching algorithms, or machine learning models. For example, if the log data contains multiple "OutOfMemoryError" occurrences, the fault category of the root cause service node can be determined to be a memory leak. Based on the fault category of the root cause service node, a pre-set table of candidate fault categories and fault repair strategies is searched to find the target repair strategy that matches the current fault category. For example, if the fault repair strategy for memory leaks in the candidate fault category-fault repair strategy mapping is to increase the Java Virtual Machine (JVM) heap memory size and reduce object creation, then increasing the JVM heap memory size and reducing object creation will be selected as the target repair strategy for the fault category. Furthermore, the target repair strategy can be executed through a configuration management tool or application programming interface (API), such as by calling the Kubernetes API to update the service deployment configuration or modify configuration center parameters. After implementing the repair, continue to monitor the health of the system to ensure that the fault has been completely resolved and adjust the subsequent operation plan based on the actual effect.

[0120] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the root cause location method of the fault of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.

[0121] This application also provides a fault root cause location device, please refer to Figure 4 , the root cause location device of the fault includes:

[0122] The pool full signal identification module 401 is used to receive the thread pool full signal sent by the message middleware and enter the fault root cause location phase;

[0123] A link graph construction module 402 is configured to construct a service node call relationship link graph based on at least one cross-service node single-step request timeout data generated in the fault root cause location phase;

[0124] The root cause location module 403 is used to determine the service node that is the root cause of the fault according to the service node call relationship link diagram.

[0125] The fault root cause locating device provided in this application, employing the fault root cause locating method described in the aforementioned embodiments, can resolve the technical problem in the prior art of being unable to accurately locate the service node at the root of a distributed system resource contention fault. Compared to the prior art, the beneficial effects of the fault root cause locating device provided in this application are the same as those of the fault root cause locating method described in the aforementioned embodiments. Other technical features of the fault root cause locating device are the same as those disclosed in the aforementioned embodiments and are not further elaborated here.

[0126] The present application provides a device for locating the root cause of a fault, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method for locating the root cause of the fault in the first embodiment described above.

[0127] Reference below Figure 5 , which shows a schematic structural diagram of a fault root cause location device suitable for implementing an embodiment of the present application. The fault root cause location device in the embodiment of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (Portable Application Descriptions, PADs), portable multimedia players (PMPs), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5The fault root cause location device shown is only an example and should not limit the functions and scope of use of the embodiments of the present application.

[0128] like Figure 5 As shown, the fault root cause locating device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. Various programs and data required for the operation of the fault root cause locating device are also stored in the RAM 1004. The processing device 1001, the read-only memory 1002, and the RAM 1004 are connected to each other via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the fault root cause location device to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows a fault root cause location device with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems may be implemented or have alternatively.

[0129] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are performed.

[0130] The fault root cause locating device provided in this application utilizes the fault root cause locating method described in the aforementioned embodiment to address the technical problem of fault root cause locating. Compared to the prior art, the beneficial effects of the fault root cause locating device provided in this application are the same as those of the fault root cause locating method described in the aforementioned embodiment. Other technical features of the fault root cause locating device are the same as those disclosed in the aforementioned embodiment and are not further elaborated here.

[0131] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0132] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0133] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, a computer program) stored thereon, wherein the computer-readable program instructions are used to execute the method for locating the root cause of a fault in the above embodiment.

[0134] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), an optical fiber, a portable compact disk read-only memory (CD-Read Only Memory, CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0135] The computer-readable storage medium may be included in the fault root cause locating device, or may exist independently without being incorporated into the fault root cause locating device.

[0136] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the root cause locating device of the fault, the root cause locating device of the fault: receives a thread pool full signal sent by the message middleware and enters the root cause locating phase of the fault; constructs a service node call relationship link graph based on at least one cross-service node single-step request timeout data generated in the root cause locating phase of the fault; and determines the root cause service node of the fault based on the service node call relationship link graph.

[0137] The computer program code for performing the operations of the present application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., using an Internet service provider to connect via the Internet).

[0138] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.

[0139] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.

[0140] The computer-readable storage medium provided in this application stores computer-readable program instructions (i.e., a computer program) for executing the aforementioned method for locating the root cause of a fault. This computer-readable storage medium can address the prior art's inability to accurately locate the service node that is the source of a distributed system resource contention fault. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the method for locating the root cause of a fault provided in the aforementioned embodiment, and are not further elaborated here.

[0141] The present application also provides a computer program product, including a computer program, which implements the steps of the above-mentioned method for locating the root cause of a fault when the computer program is executed by a processor.

[0142] The computer program product provided in this application can resolve the technical problem of the prior art in being unable to accurately locate the service node at the root of a distributed system resource contention failure. Compared to the prior art, the beneficial effects of the computer program product provided in this application are the same as those of the root cause location method provided in the aforementioned embodiment, and are not further elaborated here.

[0143] The above are only some embodiments of the present application and are not intended to limit the patent scope of the present application. All equivalent structural transformations made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A method for locating the root cause of a fault, characterized in that: Applied to a fault root cause location system, the fault root cause location method includes: Receive the thread pool full signal sent by the message middleware and enter the fault root cause identification phase; Constructing a service node call relationship link graph based on at least one cross-service node single-step request timeout data generated in the fault root cause locating phase; The service node that is the root cause of the fault is determined according to the service node call relationship link diagram.

2. The method for locating the root cause of a fault according to claim 1, wherein: Before the step of constructing a service node call relationship link graph based on at least one cross-service node single-step request timeout data generated in the fault root cause locating stage, the method further includes: Receive the connection pool full signal sent by the message middleware and enter the fault root cause location phase.

3. The method for locating the root cause of a fault according to claim 1, wherein: After entering the fault root cause location phase, the following steps are also included: At the end of each cycle in the fault root cause location phase, executing the steps of constructing a service node call relationship link graph based on at least one cross-service node single-step request timeout data generated in the fault root cause location phase; and determining the fault root cause service node based on the service node call relationship link graph; If the thread pool full signal is not detected within a preset time, the root cause location phase ends.

4. The method for locating the root cause of a fault according to any one of claims 1 or 3, wherein: The step of constructing a service node call relationship link graph based on at least one cross-service node request timeout data generated in the fault root cause location phase includes: Extracting the request identification information, single-step request identification information, requesting service node identification information, requested service node identification information, single-step request response time, single-step request response timeout time, and single-step request result from each cross-service node single-step request timeout data; Aggregating cross-service node single-step request timeout data belonging to the same request based on each of the request identification information; For each request, reconstruct the call path between each service node based on each single-step request identification information, each requesting service node identification information, and each requested service node identification information to generate an initial call tree with a time sequence relationship; Obtaining a service node initial call relationship link graph based on at least one generated initial call tree, wherein the service node initial call relationship link graph includes a plurality of vertices and a plurality of edges connecting the vertices, wherein the vertices are used to represent the service nodes, and the edges are used to represent the call relationship between two service nodes connected by the edges; Determine the timeout propagation weight of each edge based on each single-step request response time and each single-step request response timeout; The service node call relationship link graph is determined according to the timeout propagation weight of each edge and the service node initial call relationship link graph.

5. The method for locating the root cause of a fault according to claim 4, wherein: The step of determining the timeout propagation weight of each edge based on each single-step request response time and each single-step request response timeout time includes: For each timeout event, determining a timeout ratio corresponding to the edge in the timeout event according to the single-step request response time and the single-step request response timeout time; At least one of the timeout ratios of the edge is summed to obtain a timeout propagation weight of the edge.

6. The method for locating the root cause of a fault according to claim 4, wherein: The step of determining the service node that is the root cause of the fault according to the service node call relationship link diagram includes: Using a graph algorithm to perform centrality analysis on each of the service nodes in the service node call relationship link graph to obtain an importance score for each service node; Based on the importance scores of the service nodes, the service node corresponding to the service node with the largest importance score is determined as the fault root cause service node.

7. The method for locating the root cause of a fault according to claim 1, wherein: After the step of determining the service node that is the root cause of the fault according to the service node call relationship link diagram, the method further includes: Obtaining log data of the fault root cause service node, and determining a fault category of the fault root cause service node according to the log data; Determining a target repair strategy corresponding to the fault category based on a correspondence between the fault category and pre-set candidate fault categories and fault repair strategies; The target repair strategy is invoked to repair the fault of the root cause service node.

8. A method for locating the root cause of a fault, characterized in that: Applied to distributed systems, the root cause location method of the fault includes: Monitoring the resource status of the worker thread pool in each service node, and generating a thread pool full signal if the worker thread pool is determined to be full based on the resource status of the worker thread pool; monitoring the single-step request processing results of each of the service nodes, and generating cross-service node single-step request timeout data according to the single-step request processing results when detecting that the single-step request response time in the single-step request processing results is greater than the corresponding time threshold; The thread pool full signal and the cross-service node single-step request timeout data are sent to the message middleware.

9. A device for locating the root cause of a fault, characterized in that: The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the method for locating the root cause of a fault according to any one of claims 1 to 8.

10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the fault root cause locating method according to any one of claims 1 to 8 are implemented.