Fault attribution method and device, equipment and storage medium

By building a directed ring-free topology diagram and using preset detection models to calculate the fault score, the problems of slow fault positioning, low detection rate and high false alarm rate in traditional fault positioning methods are solved, and fast and accurate fault positioning and repair suggestions are achieved, improving operation and maintenance efficiency.

CN120200892APending Publication Date: 2025-06-24BEIJING BAIJU YIXING TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510374989.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Traditional fault location methods have problems such as slow fault location, low detection rate and high false alarm rate. Especially in complex multi-service, component and dynamic topology deployment platforms, it is difficult to effectively identify the root cause of the fault.

Method used

By obtaining the multi-source data information of the current service system, a directed acyclic topology diagram is built, and based on the graph and preset detection model, the fault scores of each related node are calculated, so as to quickly locate the root cause of the failure and provide repair suggestions.

Benefits of technology

It improves the fault detection rate and reduces the false alarm rate, achieves fast and accurate fault location and repair suggestions, and improves operation and maintenance efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120200892A_ABST
    Figure CN120200892A_ABST
Patent Text Reader

Abstract

The invention relates to a fault attribution method and device, equipment and a storage medium. According to the main technical scheme, the method comprises the steps of obtaining multi-source data information of a current service system, and constructing a directed acyclic topological graph according to the multi-source data information; obtaining a current fault node, and determining a related node having a causal relationship with the current fault node according to the directed acyclic topological graph; obtaining a fault weight of each related node according to a preset detection model; obtaining a fault score of each related node according to the fault weight of each related node; and according to the fault score of each related node, determining an attribution node and generating a fault attribution repair report. According to the method, the directed acyclic topological graph can be constructed based on the multi-source data information, and the fault score of each related node can be obtained based on the directed acyclic topological graph and the preset detection model, so that the system fault root cause can be quickly and accurately positioned based on the fault score, and the effect of improving the fault detection rate and the false alarm rate can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of system failures, and particularly to a fault attribution method, apparatus, device, and storage medium. Background Art

[0002] With the wide application of cloud computing and distributed systems, as well as the continuous development of Internet services, the microservices architecture of services has been increasingly favored by major enterprises. The complexity of software systems has been continuously increasing, which also brings greater challenges to traditional operation and maintenance. System failures have become the main factor affecting service stability. There are a large number of multi-dimensional KPI indicators and complex relationships among them. It is an urgent need for operation and maintenance personnel to locate faults in the first time after a failure occurs.

[0003] Traditional fault location usually adopts the method of index monitoring plus manual analysis and judgment. For example, the running memory and CPU of the server are monitored, and an alarm is triggered when the threshold is exceeded, and then manual intervention is carried out for fault analysis and repair.

[0004] However, there are some defects in the above method. First, the fault location is slow, and it usually requires manual access to classify and locate the cause of the fault. Second, the detection rate is low. Since a fixed index monitoring algorithm is adopted, all scenarios cannot be covered, and the fault detection rate is limited. For example, the CPU index monitoring usually only adopts a static critical value. Finally, the false alarm rate is high, and it cannot effectively identify the memory CPU glitch scenario, resulting in a high false alarm rate of fault detection. Moreover, especially in a deployment platform including multiple services, components, and dynamic topologies, the correlation between system logs, monitoring indicators, and topological relationships makes fault attribution more complex. Summary of the Invention

[0005] Based on this, the present application provides a fault attribution method, apparatus, device, and storage medium. Based on multi-source data information, a directed acyclic topology graph is constructed. Based on the directed acyclic topology graph and a preset detection model, the fault scores of each relevant node are obtained. Furthermore, the root cause of the fault can be quickly located based on the fault scores and repair suggestions can be provided, achieving the effect of improving the fault detection rate and false alarm rate.

[0006] In a first aspect, a fault attribution method is provided, and the method includes:

[0007] Obtain multi-source data information of the current service system, and construct a directed acyclic topology graph according to the multi-source data information;

[0008] Obtain the current fault node, and determine relevant nodes having a causal relationship with the current fault node according to the directed acyclic topology graph;

[0009] Obtain the fault weights of each relevant node according to the preset detection model;

[0010] Obtain the fault scores of each relevant node according to the fault weights of each relevant node;

[0011] Determine the attribution node and generate a fault attribution repair report according to the fault scores of each relevant node.

[0012] According to an implementable manner in the embodiment of the present application, construct a directed acyclic topology graph according to multi-source data information, including;

[0013] Perform a preprocessing operation on the multi-source data information to obtain preprocessing information;

[0014] Obtain the historical fault information of the current service system;

[0015] Construct a directed acyclic topology graph according to the preprocessing information and the historical fault information.

[0016] According to an implementable manner in the embodiment of the present application, the preset detection model includes a preset time decay model; according to the preset detection model, obtain the fault weights of each relevant node, including:

[0017] Obtain the preset cycle value, the current time value, and the preset correlation coefficient value of each relevant node in the directed acyclic topology graph that has a causal relationship with the current fault node;

[0018] According to the preset time decay model and the preset cycle value, the current time value, and the preset correlation coefficient value of each relevant node, obtain the time decay fault weight value corresponding to each relevant node.

[0019] According to an implementable manner in the embodiment of the present application, the preset detection model further includes a preset link detection model. According to the preset detection model, obtaining the fault weights of each relevant node further includes:

[0020] Obtain the preset link position weight value, the preset link weight coefficient value, and the preset basic weight coefficient value of each relevant node in the directed acyclic topology graph that has a causal relationship with the current fault node;

[0021] According to the preset link detection model and the preset link position weight value, the preset link weight coefficient value, and the preset basic weight coefficient value of each relevant node, obtain the link fault weight value corresponding to each relevant node.

[0022] According to an implementable manner in the embodiment of the present application, the preset detection model further includes a preset state transition model. According to the preset detection model, obtaining the fault weights of each relevant node further includes:

[0023] Obtain the preset state transition times value and the preset total number of fault states value of each relevant node in the directed acyclic topology graph that has a causal relationship with the current fault node;

[0024] According to a preset state transition model, the preset state transition times value and the preset total number of fault states value of each relevant node, a state transition fault weight value corresponding to each relevant node is obtained.

[0025] According to an implementable manner in an embodiment of the present application, the fault weight includes a time decay fault weight value, a link fault weight value, and a state transition fault weight value; according to the fault weights of each relevant node, a fault score of each relevant node is obtained, including:

[0026] The time decay fault weight value, the corresponding link fault weight value, and the corresponding state transition fault weight value of each relevant node are subjected to a weighted average operation to obtain the fault score of each relevant node.

[0027] According to an implementable manner in an embodiment of the present application, according to the fault scores of each relevant node, an attribution node is determined and a fault attribution repair report is generated, including:

[0028] The fault scores of each relevant node are sorted in ascending order, and the relevant node with the highest fault score is selected as the fault attribution node;

[0029] A fault attribution repair report is generated according to the fault attribution node to instruct the staff to perform repairs.

[0030] In a second aspect, a fault attribution device is provided, and the device includes:

[0031] A construction unit, configured to obtain multi-source data information of the current service system, and construct a directed acyclic topology graph according to the multi-source data information;

[0032] A determination unit, configured to obtain the current fault node, and determine relevant nodes having a causal relationship with the current fault node according to the directed acyclic topology graph;

[0033] A weight unit, configured to obtain the fault weights of each relevant node according to a preset detection model;

[0034] A scoring unit, configured to obtain the fault scores of each relevant node according to the fault weights of each relevant node;

[0035] A repair unit, configured to determine an attribution node and generate a fault attribution repair report according to the fault scores of each relevant node.

[0036] In a third aspect, a computer device is provided, including:

[0037] At least one processor; and

[0038] A memory communicatively connected to the at least one processor; wherein,

[0039] The memory stores computer instructions executable by at least one processor, and the computer instructions are executed by at least one processor to enable the at least one processor to execute the method involved in the above first aspect.

[0040] In a fourth aspect, a computer-readable storage medium is provided, on which computer instructions are stored, characterized in that the computer instructions are used to cause a computer to execute the method involved in the above first aspect.

[0041] According to the technical content provided by the embodiments of the present application, multi-source data information of the current service system is acquired, and a directed acyclic topology graph is constructed according to the multi-source data information; the current faulty node is acquired, and according to the directed acyclic topology graph, relevant nodes having a causal relationship with the current faulty node are determined; according to a preset detection model, the fault weights of each relevant node are obtained; according to the fault weights of each relevant node, the fault scores of each relevant node are obtained; according to the fault scores of each relevant node, the attribution node is determined and a fault attribution repair report is generated. The above operations construct a directed acyclic topology graph based on multi-source data information, and obtain the fault scores of each relevant node based on the directed acyclic topology graph and the preset detection model. Furthermore, the root cause of the fault can be quickly located based on the fault scores and repair suggestions can be provided, achieving the effect of improving the fault detection rate and false alarm rate. Description of the Drawings

[0042] Figure 1 It is a schematic flowchart of a fault attribution method in an embodiment;

[0043] Figure 2 It is an example of a directed acyclic topology graph of a fault attribution method in an embodiment;

[0044] Figure 3 It is a preferred schematic flowchart of a fault attribution method in an embodiment;

[0045] Figure 4 It is a structural block diagram of a fault attribution device in an embodiment;

[0046] Figure 5 It is a schematic structural diagram of a computer device in an embodiment. Detailed Embodiments

[0047] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0048] In one or more embodiments, a fault attribution method provided by the present application can be applied to a computer device, which can include a terminal or a server. Specifically, the computer device obtains multi-source data information of the current service system, constructs a directed acyclic topology graph according to the multi-source data information; obtains the current fault node, and determines related nodes that have a causal relationship with the current fault node according to the directed acyclic topology graph; obtains the fault weights of each related node according to a preset detection model; obtains the fault scores of each related node according to the fault weights of each related node; determines the attribution node and generates a fault attribution repair report according to the fault scores of each related node. Among them, the terminal can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, etc., and the server can be an independent server or a server cluster composed of multiple servers.

[0049] In one embodiment, as Figure 1 shown, a fault attribution method is provided, including the following steps:

[0050] Step S101: Obtain multi-source data information of the current service system, and construct a directed acyclic topology graph according to the multi-source data information.

[0051] Among them, the current service system includes but is not limited to a monitoring system and a middleware system. The multi-source data information includes but is not limited to log information, monitoring metric information, resource configuration information, and upstream and downstream deployment information, etc.; the reason for obtaining the multi-source data information is to achieve cross-data domain correlation analysis and comprehensive system state perception, and avoid the problem of data islands.

[0052] Here, the computer device can obtain the log information, monitoring metric information, resource configuration information, and upstream and downstream deployment information of the current service system. Among them, the log information can be the logs printed by computer programs; the monitoring metric information can be information such as CPU / memory and business monitoring; the configured resource information can be the database information on which the current service system depends; the upstream and downstream topology information can be information such as strong and weak dependencies, primary and standby calls, etc. According to the obtained multi-source data information, the multi-source data information is recursively partitioned to construct a directed acyclic topology graph. The specific structure of the directed acyclic topology graph can be as Figure 2 shown.

[0053] Step S103: Obtain the current fault node, and determine related nodes that have a causal relationship with the current fault node according to the directed acyclic topology graph.

[0054] Among them, a directed acyclic topology graph means that it is impossible to start from a certain vertex and return to that point after passing through several edges. For example, if the failure of node A may cause node B to fail to work properly, then there is a directed edge from node A to node B, but there is no connection from B to A.

[0055] Here, the computer device obtains the current faulty node. Since the directed acyclic topology graph contains the connection relationships between nodes, it is possible to determine the relevant nodes that have a causal relationship with the current faulty node based on the directed acyclic topology graph. As Figure 2 shown, assume that node 5 is the faulty node. Then the relevant nodes that have a causal relationship with this node are node 3, node 7, node 8, and node 9. And since the directed acyclic topology graph is sensed and constructed / updated in real time, it is possible to synchronize the service dependency relationships and resource allocation situations in real time, so as to achieve the effect of improving the real-time performance and accuracy of fault attribution.

[0056] Step S105: Obtain the fault weights of each relevant node according to a preset detection model.

[0057] Here, based on the preset detection model, the fault weights of each relevant node that has a causal relationship with the current faulty node can be calculated, so as to improve the accuracy and reliability of anomaly recognition in the later stage.

[0058] Step S107: Obtain the fault scores of each relevant node according to the fault weights of each relevant node.

[0059] Here, after obtaining the fault weights of each relevant node, the fault scores of each relevant node can be calculated based on each fault weight, so as to identify potential anomalies in the system and find the root cause of the fault.

[0060] Step S109: Determine the attribution node and generate a fault attribution repair report according to the fault scores of each relevant node.

[0061] Among them, the fault attribution repair report includes but is not limited to text reports and visualization charts, etc. At the same time, relevant evidence and explanations are provided to help the operation and maintenance personnel understand the cause and process of the fault occurrence.

[0062] Here, according to the fault scores of each relevant node, the root node that causes the fault can be determined, and then a fault attribution repair report can be automatically generated. Among them, the fault attribution repair report includes targeted repair suggestions, such as configuration adjustment, resource expansion, or code optimization, etc.

[0063] It can be seen that in the embodiments of the present application, by obtaining multi-source data information of the current service system, a directed acyclic topology graph is constructed according to the multi-source data information; the current faulty node is obtained, and according to the directed acyclic topology graph, related nodes having a causal relationship with the current faulty node are determined; according to a preset detection model, the fault weights of each related node are obtained; according to the fault weights of each related node, the fault scores of each related node are obtained; and according to the fault scores of each related node, the attribution node is determined and a fault attribution repair report is generated. Through the above operations, a directed acyclic topology graph is constructed based on multi-source data information, and based on the directed acyclic topology graph and the preset detection model, the fault scores of each related node are obtained. Furthermore, the root cause of the fault can be quickly located based on the fault scores and repair suggestions can be provided, achieving the effect of improving the fault detection rate and false alarm rate.

[0064] The following describes each step in the above method flow in detail. First, in combination with the embodiments, step 101 above, that is, "construct a directed acyclic topology graph according to multi-source data information", is described in detail.

[0065] Perform a preprocessing operation on the multi-source data information to obtain preprocessing information; obtain the historical fault information of the current service system; and construct a directed acyclic topology graph according to the preprocessing information and the historical fault information.

[0066] Among them, the preprocessing information includes but is not limited to data unique identifiers.

[0067] Here, since the data structures of the multi-source data information are different, the multi-source data information can be preprocessed to obtain preprocessing information. Among them, the preprocessing operations include data cleaning, normalization, feature extraction, formatting, etc. Specifically, the multi-source data information can be converted into a unified format. For example, some time formats are timestamps, while some time formats are strings, etc., and they can all be unified according to the preset standard format to obtain preprocessing information. It should be noted that the historical fault information of the current service system can also be obtained, and the connection relationships between different nodes are determined based on the data unique identifiers in the preprocessing information and the historical fault information. Furthermore, a directed acyclic topology graph can be constructed based on the data unique identifiers and the historical fault information.

[0068] Through the above operations, by constructing a directed acyclic topology graph, the deployment topology changes of the system can be sensed in real time. When a fault occurs, causal reasoning is performed using the constructed directed acyclic topology graph to analyze the fault propagation path and influence range, ensuring the accuracy of fault attribution and achieving the effect of narrowing the possible range of fault sources.

[0069] The following describes in detail step S105 above, that is, "obtain the fault weights of each related node according to the preset detection model", in combination with the embodiments.

[0070] In an implementable manner, obtain the preset cycle values, current time values, and preset correlation coefficient values of each relevant node in the directed acyclic topology graph that has a causal relationship with the current faulty node; according to the preset time decay model and the preset cycle values, current time values, and preset correlation coefficient values of each relevant node, obtain the time decay fault weight values corresponding to each relevant node.

[0071] Here, the preset detection model includes a preset time decay model, and the expression of the preset time decay model can be represented as follows:

[0072]

[0073] Among them, W represents the time decay fault weight value; T represents the preset cycle value; t represents the current time value, which can be obtained in real time; the preset correlation coefficient value includes a time decay coefficient value and a basic weight coefficient value, c represents the time decay coefficient value; b represents the basic weight coefficient value.

[0074] Specifically, obtain the values of T, t, c, and b of each relevant node in the directed acyclic topology graph that has a causal relationship with the current faulty node, and input the values of T, t, c, and b into the preset time decay model, and the time decay fault weight values of each relevant node can be obtained.

[0075] In another implementable manner, obtain the preset link position weight value, preset link weight coefficient value, and preset basic weight coefficient value of each relevant node in the directed acyclic topology graph that has a causal relationship with the current faulty node; according to the preset link detection model and the preset link position weight value, preset link weight coefficient value, and preset basic weight coefficient value of each relevant node, obtain the link fault weight values corresponding to each relevant node.

[0076] Among them, the preset link position weight value includes a preset upstream link position weight value and a preset downstream link position weight value; the preset link weight coefficient value includes a preset upstream link weight coefficient value and a preset downstream link weight coefficient value.

[0077] Here, the preset detection model further includes a preset link detection model, and the preset link detection model includes a preset upstream link detection model and a preset downstream link detection model. The specific expressions of the preset upstream link detection model and the preset downstream link detection model can be represented as follows:

[0078]

[0079] Among them, W uprepresents the upstream link failure weight value relative to the current failure node; b represents the preset basic weight coefficient value, i.e., the weight coefficient value when the link position is 0; n represents the link position, n >= 0 and n is an integer. When n = 0, it represents the current link position, and when n > 0, it represents the upstream and downstream positions from the current link position; represents the preset upstream link position weight value; W down represents the downstream link failure weight value relative to the current failure node; represents the preset downstream link position weight value.

[0080] Specifically, obtain the values of b and n for each relevant node that has a causal relationship with the current failure node in the directed acyclic topology graph, and input the values of b and n into the preset link detection model, and the link failure weight values of each relevant node can be obtained. It should be noted that since the preset link detection model has two expressions, when the relevant node is an upstream node of the current failure node, the preset upstream link detection model is used, and the values of b and n are input into the preset upstream link detection model; when the relevant node is a downstream node of the current failure node, the preset downstream link detection model is used, and the values of b and n are input into the preset downstream link detection model. Since in the link structure of the system, the influence degrees of abnormal events at different positions on the overall system are different, the above model can assign corresponding weights to events according to their positions in the link, and the events closer to the core or key positions have higher weights.

[0081] In another implementable way, obtain the preset state transition probability values and preset state transition coefficient values of each relevant node that has a causal relationship with the current failure node in the directed acyclic topology graph; according to the preset state transition model and the preset state transition probability values and preset state transition coefficient values of each relevant node, obtain the state transition failure weight values corresponding to each relevant node.

[0082] Here, the preset detection model also includes a preset state transition model, and the expression of the preset state transition model can be represented as follows:

[0083] W P = N ij / N i

[0084] Among them, W P represents the state transition probability value, i.e., the state transition failure weight value; N ij represents the number of times of transferring from the current failure node to the relevant node, i.e., the preset state transition number value, which is preset by the staff based on historical failure information; N iThe total number of times representing the fault status of relevant nodes, i.e., the preset total number of fault status values, is also preset by the staff based on historical fault information.

[0085] Specifically, obtain the N of each relevant node in the directed acyclic topology graph that has a causal relationship with the current fault node ij and N i values, and input the N ij and N i values into the preset state transition model, and the state transition fault weight values of each relevant node can be obtained.

[0086] Through the above operations, by using multiple prediction models, calculate the fault weight values of each relevant node, so that the calculation of the scores of each relevant node in the later stage is more accurate, and then abnormal events in the system can be identified more comprehensively and accurately.

[0087] Next, a detailed description will be given of the above step S107, that is, "obtain the fault weights of each relevant node according to the preset detection model", in combination with an embodiment.

[0088] Perform a weighted average operation on the time-decaying fault weight value, the corresponding link fault weight value, and the corresponding state transition fault weight value of each relevant node to obtain the fault score of each relevant node.

[0089] Here, assume that the time-decaying fault weight value is represented by W1, the link fault weight value is represented by W2, and the state transition fault weight value is represented by W3. Then the fault score of each relevant node can be represented by W a and the specific expression can be represented as follows:

[0090] W a =(W1 + W1 + W1) / 3

[0091] Through the above operations, calculate the weight values of each relevant node through multiple preset detection models, and obtain the fault scores of each relevant node through the weight values of each relevant node, making the result of the fault score more objective and accurate, significantly shortening the fault diagnosis time, and improving the accuracy of fault attribution.

[0092] Finally, a detailed description will be given of the above step S109, that is, "determine the attribution node and generate a fault attribution repair report according to the fault scores of each relevant node", in combination with an embodiment.

[0093] Sort the fault scores of each relevant node in ascending order, select the relevant node with the highest fault score as the fault attribution node; generate a fault attribution repair report according to the fault attribution node to instruct the staff to perform repairs.

[0094] Here, the fault scores of each relevant node are sorted in ascending order, and the relevant node with the highest fault score is selected as the fault attribution node. Then, a fault attribution repair report can be generated based on the fault attribution node to instruct the staff to repair the fault, achieving the effect of improving the operation and maintenance efficiency.

[0095] Through the above operation, by using the relevant node with the highest fault score as the fault attribution node, the accuracy of fault attribution is improved, the fault diagnosis time is significantly shortened, and the fault attribution repair report can be fed back in the form of interface display to instruct the staff to repair the fault, achieving the effect of improving the operation and maintenance efficiency.

[0096] Combined with the implementation methods in the above embodiments, the following will Figure 3 give an example description of a preferred method flow provided by the embodiments of the present application. As Figure 3 shown, the method may include the following steps:

[0097] Step S201, obtain multi-source data information of the current service system.

[0098] Step S202, perform preprocessing operations on the multi-source data information to obtain preprocessed information.

[0099] Step S203, obtain historical fault information of the current service system.

[0100] Step S204, construct a directed acyclic topology graph according to the preprocessed information and the historical fault information.

[0101] Step S205, obtain the preset cycle value, the current time value, and the preset correlation coefficient value of each relevant node that has a causal relationship with the current fault node in the directed acyclic topology graph.

[0102] Step S206, according to the preset time decay model and the preset cycle value, the preset correlation coefficient value, and the current time value of each relevant node, obtain the time decay fault weight value corresponding to each relevant node.

[0103] Step S207, obtain the preset link position weight value, the preset link weight coefficient value, and the preset basic weight coefficient value of each relevant node that has a causal relationship with the current fault node in the directed acyclic topology graph.

[0104] Step S208, according to the preset link detection model and the preset link position weight value, the preset link weight coefficient value, and the preset basic weight coefficient value of each relevant node, obtain the link fault weight value corresponding to each relevant node.

[0105] Step S209: Obtain the preset state transition count values and the preset total fault state count values of each relevant node in the directed acyclic topology graph that has a causal relationship with the current fault node.

[0106] Step S210: According to the preset state transition model and the preset state transition count values and the preset total fault state count values of each relevant node, obtain the state transition fault weight values corresponding to each relevant node.

[0107] Step S211: Perform a weighted average operation on the time-decaying fault weight values, the corresponding link fault weight values, and the corresponding state transition fault weight values of each relevant node to obtain the fault scores of each relevant node.

[0108] Step S212: Sort the fault scores of each relevant node in ascending order, and select the relevant node with the highest fault score as the fault attribution node.

[0109] Step S213: Generate a fault attribution repair report based on the fault attribution node to instruct the staff to perform repairs.

[0110] It should be understood that although Figures 1 - 3 the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this application, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, Figures 1 - 3 at least a part of the steps in

[0111] Figure 4 is a schematic structural diagram of a fault attribution device provided by an embodiment of this application. As Figure 4 shown, the device may include: a construction unit 301, a determination unit 303, a weight unit 305, a scoring unit 307, and a repair unit 309. The main functions of each component module are as follows:

[0112] The construction unit 301 is configured to obtain multi-source data information of the current service system and construct a directed acyclic topology graph according to the multi-source data information;

[0113] The determination unit 303 is configured to obtain the current fault node and determine the relevant nodes that have a causal relationship with the current fault node according to the directed acyclic topology graph;

[0114] A weight unit 305 for obtaining the fault weights of each relevant node according to a preset detection model;

[0115] A scoring unit 307 for obtaining the fault scores of each relevant node according to the fault weights of each relevant node;

[0116] A repair unit 309 for determining the attribution node and generating a fault attribution repair report according to the fault scores of each relevant node.

[0117] In one embodiment, the construction unit 301 is further configured to:

[0118] Perform a preprocessing operation on the multi-source data information to obtain preprocessing information;

[0119] Obtain the historical fault information of the current service system;

[0120] Construct a directed acyclic topology graph according to the preprocessing information and the historical fault information.

[0121] In one embodiment, the preset detection model includes a preset time decay model, and the weight unit 305 is further configured to:

[0122] Obtain the preset cycle value, the current time value, and the preset correlation coefficient value of each relevant node in the directed acyclic topology graph that has a causal relationship with the current fault node;

[0123] Obtain the time decay fault weight value corresponding to each relevant node according to the preset time decay model and the preset cycle value, the preset correlation coefficient value, and the current time value of each relevant node.

[0124] In one embodiment, the preset detection model includes a preset link detection model, and the weight unit 305 is further configured to:

[0125] Obtain the preset link position weight value, the preset link weight coefficient value, and the preset basic weight coefficient value of each relevant node in the directed acyclic topology graph that has a causal relationship with the current fault node;

[0126] Obtain the link fault weight value corresponding to each relevant node according to the preset link detection model and the preset link position weight value, the preset link weight coefficient value, and the preset basic weight coefficient value of each relevant node.

[0127] In one embodiment, the preset detection model includes a preset state transition model, and the weight unit 305 is further configured to:

[0128] Obtain the preset state transition times value and the preset total number of fault states value of each relevant node in the directed acyclic topology graph that has a causal relationship with the current fault node;

[0129] Based on the preset state transition model, the preset state transition times values and the preset total number of failure state values of each relevant node, the state transition failure weight values corresponding to each relevant node are obtained.

[0130] In one embodiment, the failure weight includes a time decay failure weight value, a link failure weight value, and a state transition failure weight value. The scoring unit 307 is further configured to:

[0131] Perform a weighted average operation on the time decay failure weight values, the corresponding link failure weight values, and the corresponding state transition failure weight values of each relevant node to obtain the failure scores of each relevant node.

[0132] In one embodiment, the repair unit 309 is further configured to:

[0133] Sort the failure scores of each relevant node in ascending order, and select the relevant node with the highest failure score as the failure attribution node;

[0134] Generate a failure attribution repair report according to the failure attribution node to instruct the staff to perform repairs.

[0135] For the same and similar parts among the above embodiments, reference can be made to each other. The key points of each embodiment are the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the description of the method embodiment.

[0136] It should be noted that the embodiments of the present application may involve the use of user data. In actual applications, user-specific personal data can be used in the solutions described in this article within the scope permitted by applicable laws and regulations (such as when the user clearly consents, is effectively notified to the user, and the user clearly authorizes, etc.) in accordance with the requirements of the applicable laws and regulations of the country where it is located.

[0137] According to the embodiments of the present application, the present application also provides a computer device and a computer-readable storage medium.

[0138] As Figure 5 shown, it is a block diagram of a computer device according to an embodiment of the present application. The computer device is intended to represent various forms of digital computers or mobile devices. Among them, the digital computer may include a desktop computer, a portable computer, a workbench, a personal digital assistant, a server, a mainframe computer, and other suitable computers. The mobile device may include a tablet computer, a smart phone, a wearable device, etc.

[0139] As Figure 5As shown, device 400 includes a computing unit 401, a ROM 402, a RAM 403, a bus 404, and an input / output (I / O) interface 405. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other via the bus 404. The input / output (I / O) interface 405 is also connected to the bus 404.

[0140] The computing unit 401 can execute various processes in the method embodiments of this application according to computer instructions stored in the read-only memory (ROM) 402 or computer instructions loaded into the random access memory (RAM) 403 from the storage unit 408. The computing unit 401 can be various general and / or special processing components with processing and computing capabilities. The computing unit 401 can include, but is not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. In some embodiments, the method provided by the embodiments of this application can be implemented as a computer software program tangibly contained in a computer-readable storage medium, such as the storage unit 408.

[0141] The RAM 403 can also store various programs and data required for the operation of the device 400. Part or all of the computer program can be loaded and / or installed onto the device 400 via the ROM 402 and / or the communication unit 409.

[0142] The input unit 406, the output unit 407, the storage unit 408, and the communication unit 409 in the device 400 can be connected to the I / O interface 405. Among them, the input unit 406 can be such as a keyboard, a mouse, a touch screen, a microphone, etc.; the output unit 407 can be such as a display, a speaker, an indicator light, etc. The device 400 can exchange information, data, etc. with other devices through the communication unit 409.

[0143] It should be noted that this device can also include other components necessary for normal operation. It can also only include the components necessary to implement the solution of this application, and does not necessarily include all the components shown in the figure.

[0144] Various embodiments of the systems and technologies described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof.

[0145] The computer instructions for implementing the method of this application can be written in any combination of one or more programming languages. These computer instructions can be provided to the computing unit 401, such that when the computer instructions are executed by the computing unit 401 such as a processor, the various steps involved in the method embodiments of this application are executed.

[0146] The computer-readable storage medium provided by this application can be a tangible medium that can contain or store computer instructions for executing the various steps involved in the method embodiments of this application. The computer-readable storage medium can include, but is not limited to, storage media in the forms of electronic, magnetic, optical, electromagnetic, etc.

[0147] The above specific implementation manners do not constitute a limitation on the protection scope of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of this application shall be included within the protection scope of this application.

Claims

1. A fault attribution method, characterized in that: The method comprises: Acquire multi-source data information of the current service system, and construct a directed acyclic topology graph based on the multi-source data information; Obtaining a current faulty node, and determining, according to the directed acyclic topology graph, related nodes that have a causal relationship with the current faulty node; Obtaining the fault weight of each of the related nodes according to a preset detection model; Obtaining a fault score of each of the related nodes according to the fault weight of each of the related nodes; According to the fault scores of the relevant nodes, the attribution node is determined and a fault repair report is generated.

2. The method according to claim 1, characterized in that The step of constructing a directed acyclic topology graph according to the multi-source data information includes: Performing a preprocessing operation on the multi-source data information to obtain preprocessing information; Get historical fault information of the current service system; A directed acyclic topology graph is constructed according to the preprocessing information and the historical fault information.

3. The method according to claim 1, characterized in that The preset detection model includes a preset time decay model; and obtaining the fault weight of each of the related nodes according to the preset detection model includes: Obtaining a preset period value, a current time value, and a preset correlation coefficient value of each relevant node in the directed acyclic topology graph that has a causal relationship with the current fault node; According to the preset time decay model and the preset period value, current time value and preset correlation coefficient value of each of the relevant nodes, a time decay fault weight value corresponding to each of the relevant nodes is obtained.

4. The method according to claim 3, characterized in that The preset detection model also includes a preset link detection model, and obtaining the fault weight of each of the related nodes according to the preset detection model also includes: Obtaining a preset link position weight value, a preset link weight coefficient value, and a preset basic weight coefficient value of each relevant node in the directed acyclic topology graph that has a causal relationship with the current fault node; According to the preset link detection model and the preset link position weight value, preset link weight coefficient value and preset basic weight coefficient value of each relevant node, a link fault weight value corresponding to each relevant node is obtained.

5. The method according to claim 4, characterized in that The preset detection model also includes a preset state transition model, and obtaining the fault weight of each of the related nodes according to the preset detection model also includes: Obtaining a preset state transition count value and a preset total fault state count value of each relevant node in a directed acyclic topology graph that has a causal relationship with the current fault node; According to the preset state transfer model and the preset state transfer times and the preset total fault state times of each relevant node, the state transfer fault weight value corresponding to each relevant node is obtained.

6. The method according to claim 5, characterized in that The fault weights include a time decay fault weight value, a link fault weight value, and a state transition fault weight value; and obtaining a fault score of each of the related nodes according to the fault weights of each of the related nodes includes: A weighted average operation is performed on the time decay fault weight value, the corresponding link fault weight value and the corresponding state transition fault weight value of each of the related nodes to obtain a fault score of each of the related nodes.

7. The method according to any one of claims 1 to 6, characterized in that: Determining the attribution node and generating a fault attribution repair report according to the fault score of each of the related nodes includes: Sorting the fault scores of the relevant nodes in order from low to high, and selecting the relevant node with the highest fault score as the fault attribution node; A fault repair report is generated according to the fault attribution node to instruct the staff to perform repairs.

8. A fault attribution device, characterized in that: The method comprises: A construction unit, used to obtain multi-source data information of the current service system, and construct a directed acyclic topology graph according to the multi-source data information; A determination unit, configured to obtain a current faulty node and determine, based on the directed acyclic topology graph, related nodes that have a causal relationship with the current faulty node; A weight unit, used to obtain the fault weight of each of the related nodes according to a preset detection model; A scoring unit, configured to obtain a fault score of each of the related nodes according to the fault weight of each of the related nodes; The repair unit is used to determine the attribution node and generate a fault attribution repair report according to the fault score of each of the related nodes.

9. A computer device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores computer instructions that can be executed by the at least one processor, and the computer instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having computer instructions stored thereon, characterized in that: The computer instructions are used to cause a computer to execute the method according to any one of claims 1 to 7.