Link Failure Detection Method, Apparatus, Device, and Medium

By automatically screening and detecting link sets in the microservice architecture, directed link maps are generated, which solves the problem of inefficient fault detection in the microservice architecture, and achieves efficient and accurate fault identification and propagation range display.

CN116155688BActive Publication Date: 2025-07-18BEIJING VOLCANO ENGINE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211500900.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-28
Publication Date
2025-07-18
Estimated Expiration
2042-11-28

AI Technical Summary

Technical Problem

In the prior art, microservice architecture fault detection relies on manual analysis, resulting in inefficiency and the inability to quickly identify the source and scope of the failure.

Method used

By obtaining the analysis objects and analysis types, we automatically filter the link set and perform fault detection, generate directed link maps, and intuitively display the fault propagation range.

Benefits of technology

Improves the efficiency and accuracy of fault detection, reduces detection time, and improves the accuracy and user experience of fault detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116155688B_ABST
    Figure CN116155688B_ABST
Patent Text Reader

Abstract

The present disclosure provides a link fault detection method, apparatus, device, and medium. The method includes: in response to a link detection request for a target fault exposure point, obtaining an analysis object corresponding to the target fault exposure point and the analysis type of the analysis object, and based on the analysis object and the analysis type, screening a set of links to be analyzed from multiple links, where the links are used to represent the call relationships between microservices generated when the microservices process tasks, and based on the analysis object and the analysis type, performing link fault detection on the set of links to obtain a link fault detection result of the target fault exposure point. In the embodiments of the present disclosure, by means of the analysis object and the analysis type, the links are detected, and there is no need for manual investigation of the fault exposure points one by one, which is beneficial to saving the fault detection time and thus improving the detection efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of microservices, and in particular, to a link fault detection method, apparatus, electronic device, and storage medium. Background Art

[0002] The microservice architecture is a new software architecture that realizes the functions of some complex services by deploying several microservices. Since each microservice depends on and affects each other, when any microservice fails, it will cause other microservices associated with this microservice to also have anomalies. Therefore, it is necessary to detect the source of the fault and other affected microservices. In the related art, when a microservice fails, the fault information of each microservice is usually checked manually. This method consumes a lot of manpower and takes a long time to check, resulting in low efficiency of fault troubleshooting. Summary of the Invention

[0003] Embodiments of the present disclosure at least provide a link fault detection method, apparatus, electronic device, and storage medium, which perform automated fault detection on multiple links based on the analysis object and the analysis type of the analysis object, and are beneficial to improving the efficiency of link fault detection.

[0004] Embodiments of the present disclosure provide a link fault detection method, including:

[0005] In response to a link detection request for a target fault exposure point, obtain an analysis object corresponding to the target fault exposure point and the analysis type of the analysis object; the analysis object is used to represent a target microservice corresponding to the target fault exposure point, and the analysis type is used to represent the call type of the target microservice;

[0006] Based on the analysis object and the analysis type, screen a set of links to be analyzed from multiple links; the link is used to represent the call relationship between microservices generated when the microservice processes a task;

[0007] Based on the analysis object and the analysis type, perform link fault detection on the set of links to obtain a link fault detection result of the target fault exposure point.

[0008] In the embodiments of the present disclosure, automated link fault detection is performed on the set of links based on the analysis object and the analysis type. Compared with the method of manually checking fault information in the related art, it can save detection time, thereby improving detection efficiency and detection accuracy. Further, based on the analysis object and the analysis type, a set of links to be analyzed is screened from multiple links. In this way, the number of links to be detected can be reduced in the subsequent detection step, that is, the detection range is reduced, which is beneficial to improving the detection efficiency.

[0009] In an alternative embodiment, the link failure detection of the link set based on the analysis object and the analysis type to obtain the link failure detection result of the target failure exposure point includes:

[0010] Determine the analysis starting point of the link set based on the analysis object and the analysis type; the analysis starting point is used to determine a sub-link representing the minimum failure propagation range from the link set;

[0011] According to the analysis starting point, delete the branches in the link set that are irrelevant to the analysis starting point to obtain a sub-link of the minimum failure propagation range associated with the analysis starting point;

[0012] Perform link failure detection on the sub-link of the minimum failure propagation range associated with the analysis starting point to obtain the link failure detection result of the target failure exposure point.

[0013] In the embodiments of the present disclosure, according to the analysis starting point, a sub-link of the minimum failure propagation range associated with the analysis starting point is determined from the link set, that is, the branches of each link in the link set are deleted. In this way, in the subsequent link failure detection steps, it is possible to avoid detecting branches that are irrelevant to the analysis starting point, which is beneficial to improving the efficiency of failure detection.

[0014] In an alternative embodiment, the performing link failure detection on the sub-link of the minimum failure propagation range associated with the analysis starting point in the link set to obtain the link failure detection result of the target failure exposure point includes:

[0015] Generate a directed link graph centered on the analysis starting point based on the sub-link of the minimum failure propagation range associated with the analysis starting point in the link set, and each node in the directed link graph corresponds to a microservice;

[0016] Perform failure analysis on each node in the directed link graph to obtain the link failure detection result of the target failure exposure point.

[0017] In the embodiments of the present disclosure, a directed link graph centered on the analysis starting point is generated based on the sub-link associated with the analysis starting point, and each node in the directed link graph is analyzed. In this way, the source of the failure and the impact surface of the failure can be intuitively displayed through the directed link graph, improving the user experience.

[0018] In an alternative embodiment, the fault detection result includes at least one of the number of errors, error rate, and contribution rate of each node, where the number of errors refers to the number of links with errors in each sub-link of the node, the error rate refers to the ratio of the number of errors to the total number of sub-links to which the node belongs, and the contribution rate refers to the ratio of the number of errors of the node to the total number of sub-links to which the analysis starting point belongs.

[0019] In the embodiments of the present disclosure, by determining at least one of the number of errors, error rate, and contribution rate of each node, fault detection of each node is realized from multiple detection perspectives, which is beneficial to improving the accuracy of fault detection.

[0020] In an alternative embodiment, the analysis type includes a self-call type; based on the analysis object and the analysis type, determining the analysis starting point of the link set includes:

[0021] When the analysis type is the self-call type, the analysis object is determined as the analysis starting point.

[0022] In the embodiments of the present disclosure, when the analysis type is the self-call type, it indicates that there is a fault in the analysis object itself. Therefore, the analysis object can be determined as the analysis starting point. In this way, the accuracy of the analysis starting point can be improved.

[0023] In an alternative embodiment, the link includes call information between microservices, and the analysis type includes an active call type; based on the analysis object and the analysis type, determining the analysis starting point of the link set includes:

[0024] When the analysis type is the active call type, according to the analysis object and the active call type, active call information is determined from the link set;

[0025] The calling microservice corresponding to the active call information is used as the analysis starting point of the link set.

[0026] In the embodiments of the present disclosure, when the analysis type is the active call type, first, according to the analysis object and the analysis type, active call information is determined, and the calling microservice corresponding to the active call information is determined as the analysis starting point. In this way, it is beneficial to improve the accuracy of the analysis starting point.

[0027] In an alternative embodiment, the link includes call information between microservices, and the analysis type includes a passive call type; based on the analysis object and the analysis type, determining the analysis starting point of the link set includes:

[0028] In the case where the analysis type is the passive call type, determine the passive call information from the link set according to the analysis object and the passive call type;

[0029] Use the called microservice corresponding to the passive call information as the analysis starting point of the link set.

[0030] In the embodiments of the present disclosure, in the case where the analysis type is the passive call type, first determine the passive call information according to the analysis object and the analysis type, and determine the called microservice corresponding to the passive call information as the analysis starting point. In this way, it is beneficial to improve the accuracy of the analysis starting point.

[0031] The embodiments of the present disclosure also provide a link fault detection device, including:

[0032] An acquisition module, configured to acquire an analysis object corresponding to the target fault exposure point and an analysis type of the analysis object in response to a link detection request for the target fault exposure point; the analysis object is used to represent a target microservice corresponding to the target fault exposure point, and the analysis type is used to represent a call type of the target microservice;

[0033] A screening module, configured to screen a link set to be analyzed from multiple links based on the analysis object and the analysis type; the link is used to represent a call relationship between microservices generated when the microservice processes a task;

[0034] A detection module, configured to perform link fault detection on the link set based on the analysis object and the analysis type to obtain a link fault detection result of the target fault exposure point.

[0035] In an optional implementation manner, the detection module is specifically configured to:

[0036] Determine an analysis starting point of the link set based on the analysis object and the analysis type;

[0037] The analysis starting point is used to determine a sub-link representing the minimum fault propagation range from the link set;

[0038] Delete the branches in the link set that are irrelevant to the analysis starting point according to the analysis starting point to obtain a sub-link of the minimum fault propagation range associated with the analysis starting point;

[0039] Perform link fault detection on the sub-link of the minimum fault propagation range associated with the analysis starting point to obtain a link fault detection result of the target fault exposure point.

[0040] In an optional implementation manner, the detection module is specifically configured to:

[0041] Generate a directed link graph centered on the analysis starting point based on the sub-links within the minimum fault propagation range associated with the analysis starting point in the link set, where each node in the directed link graph corresponds to a microservice;

[0042] Perform fault analysis on each node in the directed link graph to obtain the link fault detection result of the target fault exposure point.

[0043] In an alternative implementation, the fault detection result includes at least one of the number of errors, error rate, and contribution rate of each node. Among them, the number of errors refers to the number of links with errors in each sub-link of the node, the error rate refers to the ratio between the number of errors and the total number of sub-links to which the node belongs, and the contribution rate refers to the ratio between the number of errors of the node and the total number of sub-links to which the analysis starting point belongs.

[0044] In an alternative implementation, the analysis type includes the self-call type; specifically, the detection module is configured to:

[0045] When the analysis type is the self-call type, determine the analysis object as the analysis starting point.

[0046] In an alternative implementation, the link includes call information between microservices, and the analysis type includes the active call type; specifically, the detection module is configured to:

[0047] When the analysis type is the active call type, determine the active call information from the link set according to the analysis object and the active call type;

[0048] Use the calling microservice corresponding to the active call information as the analysis starting point of the link set.

[0049] In an alternative implementation, the link includes call information between microservices, and the analysis type includes the passive call type; specifically, the detection module is configured to:

[0050] When the analysis type is the passive call type, determine the passive call information from the link set according to the analysis object and the passive call type;

[0051] Use the called microservice corresponding to the passive call information as the analysis starting point of the link set.

[0052] An embodiment of the present disclosure further provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the link failure detection method in any of the above possible embodiments are executed.

[0053] An embodiment of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the steps of the link failure detection method in any of the above possible embodiments are executed.

[0054] To make the above objects, features, and advantages of the present disclosure more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, makes the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required to be used in the embodiments. The accompanying drawings are incorporated into the specification and constitute a part of this specification. These drawings show embodiments that conform to the present disclosure and, together with the specification, are used to illustrate the technical solutions of the present disclosure. It should be understood that the following drawings only show some embodiments of the present disclosure and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0056] Figure 1 Shows a flowchart of a link failure detection method provided by an embodiment of the present disclosure;

[0057] Figure 2 Shows a schematic diagram of a link provided by an embodiment of the present disclosure;

[0058] Figure 3 Shows a flowchart of a method for detecting failures in a link set provided by an embodiment of the present disclosure;

[0059] Figure 4 Shows a schematic diagram of determining an analysis starting point in an active call type provided by an embodiment of the present disclosure;

[0060] Figure 5 Shows a schematic diagram of determining an analysis starting point in a passive call type provided by an embodiment of the present disclosure;

[0061] Figure 6 Shows a flowchart of a method for detecting failures in sub-links in a link set provided by an embodiment of the present disclosure;

[0062] Figure 7 It shows a schematic diagram of a directed link graph provided by an embodiment of the present disclosure;

[0063] Figure 8 It shows a schematic diagram of a link fault detection device provided by an embodiment of the present disclosure;

[0064] Figure 9 It shows a schematic diagram of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners

[0065] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are only some of the embodiments of the present disclosure, rather than all the embodiments. Usually, the components of the embodiments of the present disclosure described and illustrated in the accompanying drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure provided in the accompanying drawings is not intended to limit the scope of the present disclosure claimed, but merely represents selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present disclosure.

[0066] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0067] The term "and / or" in this article merely describes an association relationship, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the term "at least one" in this article means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C may represent including any one or more elements selected from the set composed of A, B, and C.

[0068] The microservices architecture is an emerging software architecture. This architecture focuses on small functional blocks with single responsibilities and functions. It splits a single application program and service into several microservices and realizes corresponding functions through mutual calls between the microservices.

[0069] Since the microservices depend on and influence each other, when any one microservice fails, it is usually necessary to detect the source of its failure and its impact on other microservices. For example, the abnormal failure source of any one microservice may be propagated from other microservices several hops away, and the abnormal microservice will also affect other microservices.

[0070] In the related art, when a fault occurs, the log information of each microservice is usually checked by manual analysis. For example, manual checks are performed for each fault exposure point, and users of upstream and downstream microservices are pulled in for assistance based on the information found during the check. However, although this method can detect faults, due to the heavy dependence on manual labor and the long time consumed in the check, the efficiency of fault troubleshooting is low.

[0071] Based on the above research, the present disclosure provides a link fault detection method, apparatus, device, and storage medium. In response to a link detection request for a target fault exposure point, an analysis object corresponding to the target fault exposure point and an analysis type of the analysis object are obtained, where the analysis object is used to represent a target microservice corresponding to the target fault exposure point, and the analysis type is used to represent the call type of the target microservice. Then, based on the analysis object and the analysis type, a set of links to be analyzed is screened from multiple links, where the link is used to represent the call relationship between microservices generated when the microservice processes a task. Finally, based on the analysis object and the analysis type, link fault detection is performed on the set of links to obtain a link fault detection result for the target fault exposure point.

[0072] In an embodiment of the present disclosure, based on the analysis object and the analysis type, automated link fault detection is performed on the set of links. Compared with the method of manually checking fault information in the related art, the fault detection time can be saved, thereby improving the fault detection efficiency and accuracy.

[0073] Furthermore, in this embodiment, a set of links to be analyzed is also screened from multiple links based on the analysis object and the analysis type. In this way, the number of links to be detected can be reduced in the subsequent detection step, that is, the detection scope is narrowed, which is beneficial to improving the detection efficiency.

[0074] To facilitate the understanding of this embodiment, first, a link fault detection method disclosed in an embodiment of the present disclosure is introduced in detail. The execution subject of the link fault detection method provided in the embodiment of the present disclosure is generally an electronic device with certain computing capabilities. Such an electronic device includes, for example: a server or other processing devices. Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud storage, big data, and artificial intelligence platforms. Other processing devices can be devices including a processor and a memory, which are not limited herein. In some possible implementation manners, the link fault detection method can be implemented by a processor invoking computer-readable instructions stored in a memory.

[0075] The above link failure detection method will be introduced below with the execution entity being the server.

[0076] Refer to Figure 1 As shown, it is a flowchart of the link failure detection method provided by an embodiment of the present disclosure. The method includes steps S101 to S103, where:

[0077] S101, in response to a link detection request for a target fault exposure point, obtain an analysis object corresponding to the target fault exposure point and an analysis type of the analysis object; the analysis object is used to characterize a target microservice corresponding to the target fault exposure point, and the analysis type is used to characterize a call type of the target microservice.

[0078] Among them, a trace (also known as a call chain) refers to the tracking of a complete call of a task processing request. Specifically, in a task processing, the first microservice that executes the task processing request in a microservice architecture generates a globally unique trace identifier, which is used to identify this task processing request. During the call process of this task processing request, no matter how many microservices it passes through, this trace identifier remains unchanged and is continuously passed as each layer of microservice is called. Finally, the path of this task processing request in the microservice architecture can be strung together through the trace identifier. For example, when executing a certain request, microservice A calls microservice B, and microservice B calls microservice C and microservice D. From microservice A to microservice B, and from microservice B to microservice C and microservice D is a call chain. That is, a task processing request corresponds to a link, and each link includes the respective microservices and the call information between different microservices.

[0079] Exemplarily, please refer to Figure 2 , which is a schematic diagram of a link provided by an embodiment of the present disclosure. As Figure 2As shown in the figure, the link includes microservice S1, microservice S2, and microservice S3. Among them, M1 is the interface of microservice S1, M2 is the interface of microservice 2, and M3 is the interface of microservice 3. When S1 calls the interface M2 of S2 through interface M1, corresponding call information is generated. When S2 calls the interface M3 of S3 through interface M2, corresponding call information is generated. S2::InnerFunc refers to the internal function of S2. Here, during the microservice call process, the concepts of Server Span and Client Span are generated. For example, when S2 calls S3, S2 is the Server Span, and the process of S2 calling S3 is the Client Span (i.e., the call information). Each Span can record some information, such as which upstream microservices call S2 and which downstream microservices S2 calls, as well as information such as the response time consumed by S2 to respond to the request.

[0080] The target fault exposure point refers to the node with a fault exposed in the link. The link detection request can be a detection request generated by a user's link detection operation for the target fault exposure point.

[0081] The analysis object is used to represent the target microservice corresponding to the target fault exposure point, and the analysis type is used to represent the call type of the target microservice. The call type of the target microservice can be an active call type, a passive call type, or a self-call type.

[0082] Specifically, the user can, based on the target fault exposure point, input the analysis object corresponding to the target fault exposure point and the analysis type corresponding to the analysis object. For example, the user can input the target microservice (such as microservice S2) and the corresponding call type (such as the active call type) in a preset fault detection interface. In this way, a link detection request for the target fault exposure point can be generated and sent to the server. Then, the server can, in response to the link detection request, obtain the analysis object corresponding to the target fault exposure point and the analysis type of the analysis object. That is, in this embodiment, the analysis object and the analysis type are determined by the user; in other embodiments, the analysis object and the analysis type can also be determined by the server according to the situation of the fault exposure point. Specifically, the server can parse each fault exposure point and determine the analysis object and the analysis type according to the parsing result. For example, for the interface call information of each microservice, it is judged whether the fault occurs in the interface itself or during the interface call process, and then the analysis object and the analysis type are determined.

[0083] S102. Based on the analysis object and the analysis type, screen a set of links to be analyzed from multiple links; the link is used to represent the call relationship between microservices generated when the microservice processes a task.

[0084] It can be understood that each link is generated by the interfaces of each microservice calling each other during the process of responding to a task processing request. Each link includes each microservice required for task processing and the call information between different microservices.

[0085] Based on the above content, since one request corresponds to one link, and the analysis object is the target microservice corresponding to the target fault exposure point input by the user. For example, the analysis object input by the user can be in the format of <psm, method>, where psm is the target microservice and method is the interface of the target microservice. In this way, by using the analysis type input by the user as a further screening condition, the set of links to be analyzed that meet the screening conditions can be filtered out from multiple links.

[0086] Taking the target microservice S2 as an example, the set of links obtained for the above three analysis types will be described below.

[0087] (1) If the analysis object can be the target microservice S2 and the analysis type is the self - call type, then the call information corresponding to the interface of the target microservice S2 itself can be found, and then the corresponding set of links to be analyzed can be obtained.

[0088] (2) If the analysis object can be the target microservice S2 and the analysis type is the active - call type, then the call information corresponding to the call of the downstream microservice occurring in the interface of the target microservice S2 can be found, and the corresponding set of links to be analyzed can be obtained.

[0089] (3) If the analysis object can be the target microservice S2 and the analysis type is the passive - call type, then the corresponding call information recorded by the upstream microservice when the interface of the target microservice S2 is called by the upstream microservice can be found, and the corresponding set of links to be analyzed can be obtained.

[0090] In some embodiments, in response to a link detection request for a target fault exposure point, the set of links to be analyzed can be obtained through a real - time sampling detection mode. The real - time sampling detection mode refers to performing retrieval sampling from a real - time log storage. This mode supports flexible screening conditions. For example, it can detect the error propagation chain of "a certain cluster where an analysis object belongs calls a certain downstream microservice in a certain computer room", and multiple link acquisition tasks can also be established to achieve multi - path concurrency to improve the efficiency of link acquisition.

[0091] In some other embodiments, in response to a link detection request for a target fault exposure point, an offline subscription detection mode may be adopted to obtain a set of links to be analyzed. The offline subscription detection mode means that a user can subscribe to some fixed analysis objects, and set a timed offline task to perform subsequent fault detection on the links belonging to each analysis object. For example, obtain the set of links related to the analysis object within the most recent five days, perform subsequent fault detection on the set of links, obtain the detection results, and send the detection results to the user.

[0092] The above two detection modes can be determined by the user, or any one of the modes can be set as the default detection mode in advance, which is not limited herein.

[0093] In still some other embodiments, the offline subscription detection mode may also be to set a timed offline task to perform subsequent fault detection on the links belonging to each analysis object. In this mode, there is no need to respond to a link detection request, nor does the user need to select between the two modes. By obtaining all the links that have failed within a preset time (such as one week), and separately detecting each failed node in the links, the fault detection results can be obtained.

[0094] S103. Based on the analysis object and the analysis type, perform link fault detection on the set of links to obtain the link fault detection results of the target fault exposure point.

[0095] In this step, after obtaining the set of links through the above screening process, automated link fault detection can be performed on the set of links according to the analysis object and the analysis type to obtain the link fault detection results of the target fault exposure point. In this way, there is no need to manually check each fault exposure point one by one, which can save detection time and thus improve detection efficiency.

[0096] Specifically, for step S103, when performing link fault detection on the set of links based on the analysis object and the analysis type to obtain the link fault detection results of the target fault exposure point, please refer to Figure 3 , which includes the following S1031 to S1033:

[0097] S1031. Based on the analysis object and the analysis type, determine the analysis starting point of the set of links. The analysis starting point is used to determine a sub-link representing the minimum fault propagation range from the set of links.

[0098] Among them, the methods for determining the analysis starting point are different for different analysis types. Since the analysis type is used to represent the call type of the target microservice, the methods for determining the analysis starting point are introduced below based on different call types, specifically as follows:

[0099] (a) When the analysis type is the self - call type, the analysis object is determined as the analysis starting point.

[0100] (b) When the analysis type is the active - call type, according to the analysis object and the active - call type, determine the active - call information from the link set, and use the called microservice corresponding to the active - call information as the analysis starting point of the link set.

[0101] Exemplarily, please refer to Figure 4 , which is a schematic diagram of determining the analysis starting point in the active - call type provided by an embodiment of the present disclosure. As Figure 4 shown in, according to the analysis object S2 and the active - call type, the active - call information S2->S3::M3 can be found from the link set. In this case, it means that there is a fault when S2 calls the downstream S3. Therefore, to determine the impact surface of this fault, it is necessary to further determine its corresponding called microservice (i.e., the parent server span) from the link set based on S2->S3::M3, that is, S2::M2, and use S2::M2 as the analysis starting point of the link set.

[0102] In some embodiments, if the called microservice corresponding to the active - call information cannot be found from the link set, the subsequent detection steps can be continued by supplementing the called microservice. For example, if the called microservice corresponding to S2->S3::M3 cannot be found, S2 can be directly added to the link set as the analysis starting point of the link set.

[0103] (c) When the analysis type is the passive - call type, according to the analysis object and the passive - call type, determine the passive - call information from the link set, and use the called microservice corresponding to the passive - call information as the analysis starting point of the link set.

[0104] It can be understood that this step is similar to the content of step (b) above. Exemplarily, please refer to Figure 5 , which is a schematic diagram of determining the analysis starting point in the passive - call type provided by an embodiment of the present disclosure. As Figure 5As shown in , according to the analysis object and the passive call type, the passive call information S1->S2::M2 can be found from the link set, which means that there is a fault when S2 is called by the upstream S1 as the called microservice, that is, the source of the fault is determined here, and therefore, it is necessary to further determine the impact area of the fault, and it is necessary to further determine the corresponding called microservice (that is, the sub-server span) from the link set based on S1->S2::M2, that is, S2::M2, and use S2::M2 as the analysis starting point of the link set.

[0105] Optionally, when the called microservice is found, it can be determined whether the state of the called microservice is an error state. If the state of the called microservice is not an error state, its state is marked as an error state, and the reason for marking the error state is that the process called by the upstream is wrong.

[0106] In some implementations, if the calling microservice corresponding to the passive calling information is not found in the link set, the subsequent detection steps can be continued by supplementing the called microservice. For example, if the called microservice corresponding to S1->S2::M2 is not found, S2 can be directly added to the link set as the analysis starting point of the link set.

[0107] S1032: According to the analysis starting point, branches in the link set that are not related to the analysis starting point are deleted to obtain a sub-link with a minimum fault propagation range associated with the analysis starting point.

[0108] In this step, based on the analysis starting point, the branches in the link set that are not related to the analysis starting point are deleted to obtain the sub-link with the minimum fault propagation range associated with the analysis starting point. Specifically, for each link in the link set, since these links share the same analysis starting point, the branches in each link that are not related to the analysis starting point can be deleted separately.

[0109] Exemplarily, if the analysis starting point is S2 and the analysis type is an active call type, then when obtaining the link set, the relevant complete links are obtained. For example, the obtained links include S1->S2, S2->S3, S2->S4, that is, S2 calls two microservices S3 and S4 at the same time, and there is a fault in S2->S3, and there is no fault in the S2->S4 branch. Therefore, the S2->S4 branch can be deleted from the link, and only the S2->S3 branch is retained. That is, in this step, starting from the analysis starting point, irrelevant brother nodes are deleted, and only the direct upstream microservices and direct downstream microservices of the analysis starting point are retained. In this way, after all links are deleted, the sub-link associated with the analysis starting point can be obtained.

[0110] It should be noted that the number of microservices in this embodiment is only illustrative. In other embodiments, the number of microservices can be any number, which is not limited herein.

[0111] In some embodiments, the user can select the "shield sub-error links that have not propagated to the analysis starting point" mode in the fault detection interface. That is, if a fault occurs in the direct downstream microservices of the analysis starting point, but the fault has not propagated to the analysis starting point, that is, it has no impact on the analysis starting point, then the direct downstream microservices with faults can be deleted.

[0112] S1033. Perform link fault detection on the sub-links within the minimum fault propagation range associated with the analysis starting point in the link set to obtain the link fault detection result of the target fault exposure point.

[0113] It can be understood that after screening the link set in the above steps and deleting the branches in the links, the finally obtained sub-links are relatively concise links. Furthermore, link fault detection can be performed on the sub-links associated with the analysis starting point to obtain the link fault detection result of the target fault exposure point.

[0114] Optionally, for step S1033, when performing link fault detection on the sub-links associated with the analysis starting point in the link set to obtain the link fault detection result of the target fault exposure point, please refer to Figure 6 and it may include the following S10331 to S10332:

[0115] S10331. Based on the sub-links within the minimum fault propagation range associated with the analysis starting point in the link set, generate a directed link graph centered on the analysis starting point, where each node in the directed link graph corresponds to a microservice.

[0116] S10332. Perform fault analysis on each node in the directed link graph to obtain the link fault detection result of the target fault exposure point.

[0117] Exemplarily, please refer to Figure 7 which is a schematic diagram of a directed link graph provided by an embodiment of the present disclosure. As Figure 7 shown in it, each sub-link is centered on the analysis starting point and is connected to the analysis starting point. Among them, each circular node in the directed link graph represents a microservice, and the arrow between two nodes represents the call relationship between two microservices. For example, S1, S3, and S5 respectively call S2, and S2 in turn calls S4, S6, and S8.

[0118] Among them, the link fault detection result includes at least one of the number of errors, error rate, and contribution rate of each node.

[0119] Wherein, the number of errors refers to the number of links with errors in each sub-link of the node, the error rate refers to the ratio between the number of errors and the total number of sub-links to which the node belongs, and the contribution rate refers to the ratio between the number of errors of the node and the total number of sub-links to which the analysis starting point belongs.

[0120] Exemplarily, since each request corresponds to one link, for Request 1, Request 2, and Request 3, the links corresponding to them all include the microservice S1 calling the microservice S2. Among them, the call of S1 to S2 fails in the links corresponding to Request 1 and Request 2, while the call of S1 to S2 does not fail in the link corresponding to Request 3. Then, for the node S2, the number of errors is 2, the total number is 3, and the error rate is 2 / 3.

[0121] Taking Figure 7 S2 in [Example] as an example of the analysis starting point, the method for determining the contribution rate is described. The total number of sub-links to which the analysis starting point belongs is the total number of links in the link set determined in the foregoing embodiment. That is, the number of links in the link set determined based on the analysis object S2 is 100. S2 calls S4, S6, and S8 at the same time. Among them, the number of times S2 calls S4 is 95 times, the number of times S2 calls S6 is 4 times, and the number of times S2 calls S8 is 1 time. Then, for the node S4, the corresponding contribution rate is 95 / 100, for the node S6, the corresponding contribution rate is 4 / 100, and for the node S8, the corresponding contribution rate is 1 / 100.

[0122] Optionally, after the fault detection, the number of errors, error rate, contribution rate, etc. of each node can be marked in the directed link graph, and in response to the user's viewing request for the directed link graph, the directed link graph can be displayed. In this way, the user can intuitively analyze the fault information of each node.

[0123] Also optionally, after each request is executed, if a certain node fails, an error status code of the node and the key error log corresponding to the error status code can be generated. Furthermore, the error status code of the node and the key error log can be marked in the directed link graph, and the link identifier to which the node belongs can be marked for the user to view.

[0124] In some embodiments, there may be some branches without faults in the directed link graph. If all branches are displayed in the directed link graph, it will be redundant. Then, the nodes can be deleted according to the call ratio of the nodes. For example, if S2 calls S3 99 times, while S2 calls S4 only 1 time. At this time, the relevance between S4 and S2 is only 1%, then the S4 node can be deleted.

[0125] Those skilled in the art can understand that in the above methods of the specific embodiments, the writing order of each step does not mean a strict execution order and does not impose any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.

[0126] Based on the same inventive concept, an embodiment of the present disclosure also provides a link fault detection device corresponding to the link fault detection method. Since the principle of the device in the embodiment of the present disclosure to solve the problem is similar to the above link fault detection method in the embodiment of the present disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0127] Referring to Figure 8 As shown, it is a schematic diagram of a link fault detection device 800 provided by an embodiment of the present disclosure. The device includes:

[0128] An acquisition module 801, configured to acquire an analysis object corresponding to the target fault exposure point and an analysis type of the analysis object in response to a link detection request for the target fault exposure point; the analysis object is used to represent a target microservice corresponding to the target fault exposure point, and the analysis type is used to represent a call type of the target microservice;

[0129] A screening module 802, configured to screen a set of links to be analyzed from multiple links based on the analysis object and the analysis type; the links are used to represent call relationships between microservices generated when the microservice processes tasks;

[0130] A detection module 803, configured to perform link fault detection on the set of links based on the analysis object and the analysis type to obtain a link fault detection result of the target fault exposure point.

[0131] In an optional implementation manner, the detection module 803 is specifically configured to:

[0132] Determine an analysis starting point of the set of links based on the analysis object and the analysis type; the analysis starting point is used to determine a sub-link representing the minimum fault propagation range from the set of links;

[0133] Delete branches in the set of links that are irrelevant to the analysis starting point according to the analysis starting point to obtain a sub-link of the minimum fault propagation range associated with the analysis starting point;

[0134] Perform link fault detection on the sub-link of the minimum fault propagation range associated with the analysis starting point to obtain a link fault detection result of the target fault exposure point.

[0135] In an optional implementation manner, the detection module 803 is specifically configured to:

[0136] Generate a directed link graph centered on the analysis starting point based on the sub-links within the minimum fault propagation range associated with the analysis starting point in the link set, where each node in the directed link graph corresponds to a microservice;

[0137] Perform fault analysis on each node in the directed link graph to obtain the link fault detection result of the target fault exposure point.

[0138] In an optional implementation manner, the fault detection result includes at least one of the number of errors, error rate, and contribution rate of each node, where the number of errors refers to the number of links with errors in each sub-link of the node, the error rate refers to the ratio between the number of errors and the total number of sub-links to which the node belongs, and the contribution rate refers to the ratio between the number of errors of the node and the total number of sub-links to which the analysis starting point belongs.

[0139] In an optional implementation manner, the analysis type includes the self-call type; the detection module 803 is specifically configured to:

[0140] When the analysis type is the self-call type, determine the analysis object as the analysis starting point.

[0141] In an optional implementation manner, the link includes call information between microservices, and the analysis type includes the active call type; the detection module 803 is specifically configured to:

[0142] When the analysis type is the active call type, determine the active call information from the link set according to the analysis object and the active call type;

[0143] Use the calling microservice corresponding to the active call information as the analysis starting point of the link set.

[0144] In an optional implementation manner, the link includes call information between microservices, and the analysis type includes the passive call type; the detection module 803 is specifically configured to:

[0145] When the analysis type is the passive call type, determine the passive call information from the link set according to the analysis object and the passive call type;

[0146] Use the called microservice corresponding to the passive call information as the analysis starting point of the link set.

[0147] For the description of the processing flow of each module in the device and the interaction flow between modules, reference may be made to the relevant descriptions in the above method embodiments, which will not be elaborated here.

[0148] Based on the same inventive concept, embodiments of the present disclosure also provide an electronic device. Referring to Figure 9 As shown, it is a schematic structural diagram of an electronic device 900 provided by an embodiment of the present disclosure, including a processor 901, a memory 902, and a bus 903. Among them, the memory 902 is used to store execution instructions, including an internal memory 9021 and an external memory 9022; the internal memory 9021 here is also called the main memory, which is used to temporarily store the operation data in the processor 901 and the data exchanged with the external memory 9022 such as a hard disk. The processor 901 exchanges data with the external memory 9022 through the internal memory 9021.

[0149] In an embodiment of the present application, the memory 902 is specifically used to store the application program code for implementing the solution of the present application, and is controlled and executed by the processor 901. That is, when the electronic device 900 runs, the processor 901 communicates with the memory 902 through the bus 903, so that the processor 901 executes the application program code stored in the memory 902, and further executes the method described in any of the foregoing embodiments.

[0150] Among them, the memory 902 may be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read - only memory (PROM), an erasable programmable read - only memory (EPROM), an electrically erasable programmable read - only memory (EEPROM), etc.

[0151] The processor 901 may be an integrated circuit chip with signal processing capabilities. The above - mentioned processor may be a general - purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application - specific integrated circuit (ASIC), a field - programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general - purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0152] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the electronic device 900. In other embodiments of the present application, the electronic device 900 may include more or fewer components than those illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0153] Embodiments of the present disclosure also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of link failure detection in the above method embodiments. Among them, the storage medium may be a volatile or non-volatile computer-readable storage medium.

[0154] Embodiments of the present disclosure also provide a computer program product, which carries program code. The instructions included in the program code can be used to execute the steps of link failure detection in the above method embodiments. For details, reference can be made to the above method embodiments, which will not be elaborated here.

[0155] Among them, the above computer program product can be specifically implemented in the form of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is specifically embodied as a computer storage medium. In another alternative embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.

[0156] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems and devices can refer to the corresponding processes in the foregoing method embodiments, which will not be elaborated here. In several embodiments provided by the present disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings, direct couplings, or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0157] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0158] In addition, in each embodiment of the present disclosure, each functional unit can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0159] If the above-mentioned function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present disclosure. The aforementioned storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories, random access memories, magnetic disks, or optical discs.

[0160] Finally, it should be noted that: the above-mentioned embodiments are only specific implementation manners of the present disclosure, used to illustrate the technical solutions of the present disclosure, rather than limiting them. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present disclosure can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be covered by the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.

Claims

1. A link fault detection method, characterized in that, Including: In response to a link detection request for a target fault exposure point, obtaining an analysis object corresponding to the target fault exposure point and an analysis type of the analysis object; The analysis object is used to characterize a target microservice corresponding to the target fault exposure point, and the analysis type is used to characterize a call type of the target microservice; Based on the analysis object and the analysis type, screening a set of links to be analyzed from multiple links; the link refers to a trace of a complete call of a task processing request, and the link includes each microservice that executes the task processing request and call information between different microservices, and is used to characterize the call relationship between microservices generated when the microservice processes a task; Based on the analysis object and the analysis type, performing link fault detection on the set of links to obtain a link fault detection result of the target fault exposure point; wherein the target fault exposure point is a node with a fault in the link; wherein the performing link fault detection on the set of links based on the analysis object and the analysis type to obtain a link fault detection result of the target fault exposure point includes: Based on the analysis object and the analysis type, determining an analysis starting point of the set of links; the analysis starting point is used to determine a sub-link representing the minimum fault propagation range from the set of links; According to the analysis starting point, deleting branches in the set of links that are irrelevant to the analysis starting point to obtain a sub-link of the minimum fault propagation range associated with the analysis starting point; Performing link fault detection on the sub-link of the minimum fault propagation range associated with the analysis starting point to obtain a link fault detection result of the target fault exposure point.

2. The method according to claim 1, wherein The performing link fault detection on the sub-link of the minimum fault propagation range associated with the analysis starting point to obtain a link fault detection result of the target fault exposure point includes: Based on the sub-link of the minimum fault propagation range associated with the analysis starting point in the set of links, generating a directed link graph centered on the analysis starting point, and each node in the directed link graph corresponds to a microservice; Performing fault analysis on each node in the directed link graph to obtain a link fault detection result of the target fault exposure point.

3. The method according to claim 2, wherein The fault detection result includes at least one of the number of errors, error rate, and contribution rate of each node, wherein the number of errors refers to the number of links with errors in each sub-link of the node, the error rate refers to the ratio between the number of errors and the total number of sub-links to which the node belongs, and the contribution rate refers to the ratio between the number of errors of the node and the total number of sub-links to which the analysis starting point belongs.

4. The method according to claim 1, wherein The determining an analysis starting point of the set of links based on the analysis object and the analysis type includes: In the case where the call type is a self-call type, determining the analysis object as the analysis starting point.

5. The method according to claim 1, wherein The link includes call information between microservices; The determining an analysis starting point of the set of links based on the analysis object and the analysis type includes: In the case where the analysis type is the active call type, determine the active call information corresponding to the analysis object from the link set according to the analysis object and the active call type; Use the calling microservice corresponding to the active call information as the analysis starting point of the link set.

6. The method according to claim 1, characterized in that, The determining the analysis starting point of the link set based on the analysis object and the analysis type includes: In the case where the call type is the passive call type, determine the passive call information from the link set according to the analysis object and the passive call type; Use the called microservice corresponding to the passive call information as the analysis starting point of the link set.

7. A link fault detection device, characterized in that, It includes: An acquisition module, configured to acquire an analysis object corresponding to the target fault exposure point and an analysis type of the analysis object in response to a link detection request for the target fault exposure point; The analysis object is used to represent a target microservice corresponding to the target fault exposure point, and the analysis type is used to represent a call type of the target microservice; A screening module, configured to screen a link set to be analyzed from multiple links based on the analysis object and the analysis type; the link refers to the tracking of a complete call of a task processing request, and the link includes each microservice that executes the task processing request and call information between different microservices, and is used to represent the call relationship between microservices generated when the microservice processes the task; A detection module, configured to perform link fault detection on the link set based on the analysis object and the analysis type to obtain a link fault detection result of the target fault exposure point, wherein the target fault exposure point is a node with a fault in the link; wherein the detection module is specifically configured to: Determine the analysis starting point of the link set based on the analysis object and the analysis type; the analysis starting point is used to determine a sub-link representing the minimum fault propagation range from the link set; According to the analysis starting point, delete the branches in the link set that are irrelevant to the analysis starting point to obtain a sub-link of the minimum fault propagation range associated with the analysis starting point; Perform link fault detection on the sub-link of the minimum fault propagation range associated with the analysis starting point to obtain a link fault detection result of the target fault exposure point.

8. An electronic device, characterized in that, It includes: A processor, a memory, and a bus, the memory stores machine-readable instructions executable by the processor, when the electronic device runs, the processor communicates with the memory through the bus, and when the machine-readable instructions are executed by the processor, the steps of the link fault detection method according to any one of claims 1 to 7 are executed.

9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, the steps of the link fault detection method according to any one of claims 1 to 7 are executed.

Citation Information

Patent Citations

  • Service fault positioning method, device, computer equipment and storage medium

    CN108833184A

  • Fault equipment positioning method and device, electronic equipment, medium and program product

    CN114710400A