A fault locating method, electronic device, storage medium and program product
By using multi-dimensional feature attributes and fault tree analysis, the problems of inaccurate fault location and low efficiency in existing technologies are solved, and more efficient fault root cause identification is achieved.
Patent Information
- Application Number
- CN202511270472.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-08
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-09-08
AI Technical Summary
Existing technologies cannot fully and accurately pinpoint the root cause of a fault, and the pinpointing efficiency is low.
The similarity between the target fault and historical faults is determined by multi-dimensional feature attributes, a fault tree is constructed and the minimum cut set is extracted to determine candidate root causes.
It improves the accuracy and efficiency of fault location and avoids redundant troubleshooting.
Smart Images

Figure CN120762953B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of system operation and maintenance technology, and in particular to fault location methods, electronic devices, storage media and program products. Background Technology
[0002] Fault localization refers to the process of identifying the root cause of a fault when it is detected during the operation of a system, software, or hardware. Fault localization is a crucial step in operation and maintenance, software debugging, and system reliability engineering, directly impacting fault repair efficiency. Related technologies for fault localization typically rely on similarity matching with historical faults, selecting the historical fault with the highest similarity and reusing its fault localization results. However, these technologies cannot comprehensively and accurately locate the root cause of faults, and their localization efficiency is low. Therefore, how to overcome these technical shortcomings has become a pressing technical problem for those skilled in the art. Summary of the Invention
[0003] This invention provides a fault location method, electronic device, storage medium, and program product to at least solve the problems of incomplete and inaccurate fault location and low location efficiency in related technologies.
[0004] This invention provides a fault location method, comprising:
[0005] The similarity between the target fault and the historical fault is determined based on the multi-dimensional feature attributes of the target fault and the multi-dimensional feature attributes of the historical fault.
[0006] Select the first preset number of historical faults with the highest similarity to obtain a set of similar faults;
[0007] A fault tree is constructed based on the set of similar faults; the fault phenomena of historical faults in the set of similar faults are used as the top events of the fault tree, and the root causes of historical faults in the set of similar faults are used as the basic events of the fault tree.
[0008] Extract the minimum cut set of the fault tree; the minimum cut set includes several of the basic events;
[0009] Determine the importance of the basic events in the minimum cut set, and select the basic event with the highest importance as the candidate root cause of the target fault.
[0010] The present invention also provides a fault location device, comprising:
[0011] The similarity determination module is used to determine the similarity between the target fault and the historical fault based on the feature attributes of the target fault in multiple dimensions and the feature attributes of the historical fault in multiple dimensions.
[0012] The historical fault selection module is used to select the top preset number of historical faults with the highest similarity to obtain a set of similar faults;
[0013] A fault tree construction module is used to construct a fault tree based on the set of similar faults; the fault phenomena of historical faults in the set of similar faults are used as the top events of the fault tree, and the root causes of historical faults in the set of similar faults are used as the basic events of the fault tree.
[0014] The minimum cut set extraction module is used to extract the minimum cut set of the fault tree; the minimum cut set includes several of the basic events.
[0015] The root cause determination module is used to determine the importance of basic events in the minimum cut set and select the basic event with the highest importance as the candidate root cause of the target fault.
[0016] The present invention also provides an electronic device, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of any of the above-described fault location methods.
[0017] The present invention also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of any of the above-described fault location methods.
[0018] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the steps of any of the above-described fault location methods.
[0019] The beneficial effects are as follows: This invention determines the similarity between the target fault and historical faults based on multiple dimensions of feature attributes, which avoids the one-sidedness of similarity calculation based on a single dimension, improves the accuracy of similarity, and thus improves the accuracy of fault screening based on similarity. Furthermore, this invention constructs a fault tree corresponding to historical faults and extracts the minimum cut set to determine the candidate root causes of the target fault, thereby avoiding redundant investigation and improving fault location efficiency. Attached Figure Description
[0020] To more clearly illustrate the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a flowchart illustrating a fault location method provided in an embodiment of the present invention.
[0022] Figure 2A flowchart of a minimum cut set extraction method provided in an embodiment of the present invention;
[0023] Figure 3 A fault tree-based fault location flowchart is provided as an embodiment of the present invention;
[0024] Figure 4 A schematic diagram of a fault location device provided in an embodiment of the present invention;
[0025] Figure 5 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present invention.
[0027] It should be noted that, in the description of this invention, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., used in this invention are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0028] To enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0029] The present invention provides a fault location method. The method is described in detail below in conjunction with the execution flow of the fault location method.
[0030] refer to Figure 1 As shown, an embodiment of the present invention provides a fault location method including:
[0031] S101: Determine the similarity between the target fault and the historical fault based on the multiple dimensions of the target fault's feature attributes and the multiple dimensions of the historical fault's feature attributes.
[0032] In operating system maintenance, fault location processes and solutions are collected and organized to form a fault knowledge base. When a new fault occurs, the fault location results and troubleshooting processes of historical faults similar to the new fault can provide direction and narrow down the location scope. Therefore, this embodiment determines the similarity between the target fault and historical faults based on multiple dimensions of the target fault's characteristic attributes and those of historical faults. The target fault refers to the fault to be located.
[0033] In some embodiments, determining the similarity between the target fault and the historical fault based on multiple dimensions of feature attributes of the target fault and multiple dimensions of feature attributes of historical faults includes:
[0034] Based on the fault type, fault qualifiers, fault trigger, fault source, fault impact range, fault duration, fault detectability, and fault-related resources of the target fault, and the fault type, fault qualifiers, fault trigger, fault source, fault impact range, fault duration, fault detectability, and fault-related resources of the historical fault, the similarity between the target fault and the historical fault is determined.
[0035] In this embodiment, fault type, fault qualifier, fault trigger, fault source, fault impact range, fault duration, fault detectability, and fault-related resources are selected as fault feature attributes. Similarity calculation is performed based on the above multiple feature attributes, which can improve the accuracy of similarity and thus improve the accuracy of historical fault screening based on similarity, avoiding omissions and misselections.
[0036] Orthogonal fault classification methods categorize faults into eight attributes: fault activity, fault trigger, fault impact, fault repair object, fault type, fault qualifier, fault history, and fault source. These attributes are pairwise orthogonal and describe the characteristics of the fault from different perspectives. In the operating system domain, fault source refers to the module to which the fault belongs; fault trigger refers to the triggering conditions or actions; fault qualifier refers to the error message when the fault occurs; fault type refers to the type of fault, such as hardware or software fault. Fault activity can include the manifestation of the fault, fault impact can include the degree or result of the fault, fault repair object refers to the object that needs to be repaired, and fault history can include the number of times similar faults have occurred in the past, etc.
[0037] This invention selects four attributes—fault source, fault trigger, fault qualifier, and fault type—as characteristic attributes of the fault. Furthermore, this embodiment adds fault impact, fault duration, fault detectability, and fault-related resources as characteristic attributes to the above four attributes.
[0038] The scope of impact describes the breadth and depth of a failure's effects, reflecting the degree to which system functionality is impaired. The scope of impact can be categorized as single-user, multi-user, system cluster, and cross-platform. Single-user impact refers to affecting a single user's process or session, such as desktop lag. Multi-user impact refers to affecting multiple users within the same system, such as a shared service crash. System cluster impact refers to affecting distributed clusters or high-availability groups, such as a master node failure causing cluster failure. Cross-platform impact refers to environments involving hybrid cloud or multiple operating systems, such as cross-node communication failures in container networks.
[0039] Fault duration describes the time characteristics from the occurrence of a fault to its repair, and is related to system recovery efficiency. Fault duration can be categorized into transient faults, temporary faults, and persistent faults. Transient faults are faults that heal themselves after a brief anomaly, such as network jitter causing a temporary system outage. Temporary faults are faults that require manual intervention but can be repaired quickly, such as manual recovery after an Nginx service crash. Persistent faults are faults that require deep repair or hardware replacement, such as hard drive bad sectors or kernel-level vulnerabilities.
[0040] Fault detectability describes the ease with which a fault can be identified by a monitoring system or tool. It can be categorized into three types: proactive alarm-based, log-traceable, and silent fault-based. Proactive alarm-based faults are those that can be captured in real-time by monitoring tools, such as Zabbix triggering CPU overload alarms. Log-traceable faults are those where the application crashes and logs are retained, but no system alarms are triggered. Silent faults are those without explicit error messages, requiring inference based on business metrics, such as memory leaks causing service response delays.
[0041] Fault-related resources are used to identify the types of system resources directly associated with a fault, and can be categorized into weakly bound, CPU-bound, I / O-bound, and hybrid dependency types. Weakly bound resources are only related to the business logic of the service itself. CPU-bound resources refer to memory leaks or Out of Memory (OOM) errors, such as JVM heap overflow. I / O-bound resources refer to disk or network throughput bottlenecks, such as frequent database read / write operations causing I / O waits. Hybrid dependency types refer to concurrent contention for multiple resources, such as simultaneous CPU, memory, and network overload in high-concurrency scenarios.
[0042] Referring to Table 1, faults are quantitatively described based on fault type, fault qualifiers, fault trigger, fault source, fault impact range, fault duration, fault detectability, and fault-related resources, forming a feature set. .
[0043] in, Indicate the type of failure, such as "configuration error" or "service interruption".
[0044] These are error qualifiers, referring to the error messages displayed when a failure occurs, such as "file does not exist".
[0045] This indicates a fault trigger, referring to the action that triggers the fault, such as "restarting the service".
[0046] Indicates the source of the fault, referring to the module to which the fault belongs, such as "database service".
[0047] This indicates the scope of the fault's impact, referring to the breadth and depth of its effects and reflecting the degree of damage to system functions, such as "single-user level" or "system cluster level".
[0048] Indicates the duration of a fault, referring to the time characteristic from the occurrence of a fault to its repair, such as "transient fault" or "persistent fault".
[0049] This indicates the detectability of a fault, referring to the ease with which a fault can be identified by a monitoring system or tool, such as "active alarm type" or "silent fault type".
[0050] This indicates the resources associated with the fault, referring to the type of system resources directly associated with the fault, such as "memory-bound" or "I / O-bound".
[0051] Table 1. Meaning of Fault Characteristic Attributes
[0052]
[0053] set up For a fault set, Indicates a fault, and It can be represented as an eigenvector. ,in for In attributes The specific description above.
[0054] Based on the fault type, fault qualifiers, fault trigger, fault source, fault impact range, fault duration, fault detectability, and fault-related resources of the target fault, and compared with the fault type, fault qualifiers, fault trigger, fault source, fault impact range, fault duration, fault detectability, and fault-related resources of historical faults, the similarity between the target fault and historical faults is determined.
[0055] By describing faults in multiple dimensions, the similarity between target faults and historical faults can be determined based on these multi-dimensional feature descriptions. This effectively avoids the one-sidedness of similarity calculation based on single-dimensional features and improves the accuracy of historical fault screening.
[0056] In some embodiments, determining the similarity between the target fault and the historical fault based on multiple dimensions of feature attributes of the target fault and multiple dimensions of feature attributes of historical faults includes:
[0057] Determine the matching degree between the feature attributes of the target fault in each dimension and the feature attributes corresponding to the historical fault;
[0058] Based on the matching degree, the similarity between the target fault and the historical fault is determined.
[0059] Based on the implementation of examples where fault characteristic attributes include fault type, fault qualifiers, fault trigger, fault source, fault impact range, fault duration, fault detectability, and fault-related resources, the matching degree between the characteristic attributes of the target fault in each dimension and the corresponding characteristic attributes of historical faults is determined. This includes: determining the matching degree between the fault type of the target fault and the fault type of historical faults; determining the matching degree between the fault qualifiers of the target fault and the fault qualifiers of historical faults; determining the matching degree between the fault trigger of the target fault and the fault trigger of historical faults; determining the matching degree between the fault source of the target fault and the fault source of historical faults; determining the matching degree between the fault impact range of the target fault and the fault impact range of historical faults; determining the matching degree between the fault duration of the target fault and the fault duration of historical faults; determining the matching degree between the fault detectability of the target fault and the fault detectability of historical faults; and determining the matching degree between the fault-related resources of the target fault and the fault-related resources of historical faults. Based on the obtained matching degrees, the similarity between the target fault and historical faults is determined.
[0060] In some embodiments, determining the matching degree between the feature attributes of the target fault in each dimension and the feature attributes corresponding to the historical fault includes:
[0061] Determine the cosine similarity between the feature attributes of the target fault and the feature attributes corresponding to the historical faults;
[0062] Compare the cosine similarity with a preset threshold;
[0063] If the cosine similarity is greater than or equal to the preset threshold, then the matching degree between the feature attributes of the target fault and the feature attributes corresponding to the historical fault is equal to the cosine similarity.
[0064] If the cosine similarity is less than the preset threshold, then the matching degree between the feature attributes of the target fault and the feature attributes corresponding to the historical fault is zero.
[0065] The cosine similarity between the feature attributes of the target fault and the feature attributes corresponding to historical faults can be expressed as:
[0066] ; x represents the description of the target fault in dimension j, and y represents the description of the historical fault in dimension j. For example, x can be the description of the fault type of the target fault, and y can be the description of the fault type of the historical fault.
[0067] Based on the implementation of fault characteristic attributes including fault type, fault qualifiers, fault trigger, fault source, fault impact range, fault duration, fault detectability, and fault-related resources, the following methods are used to determine the cosine similarity between the fault type of the target fault and the fault type of historical faults; the cosine similarity between the fault qualifiers of the target fault and the fault qualifiers of historical faults; the cosine similarity between the fault trigger of the target fault and the fault trigger of historical faults; the cosine similarity between the fault source of the target fault and the fault source of historical faults; the cosine similarity between the fault impact range of the target fault and the fault impact range of historical faults; the cosine similarity between the fault duration of the target fault and the fault duration of historical faults; the cosine similarity between the fault detectability of the target fault and the fault detectability of historical faults; and the cosine similarity between the fault-related resources of the target fault and the fault-related resources of historical faults. Based on the obtained cosine similarities, the similarity between the target fault and historical faults is determined.
[0068] if If the target fault's characteristic attributes match the historical fault's characteristic attributes, the matching degree is equal to the preset threshold. .if If the threshold is less than the preset threshold, the feature attributes representing the target fault do not match the feature attributes corresponding to historical faults, and the matching degree is zero.
[0069] Taking a preset threshold of 0.8 as an example, the matching degree between each feature attribute of the target fault and the feature attributes corresponding to historical faults is expressed as follows:
[0070] . The characteristic attribute j representing the target fault, The characteristic attribute j representing historical faults, Indicates the degree of matching.
[0071] In some embodiments, the method further includes: adjusting the value of the preset threshold according to the fault type of the target fault.
[0072] For example, if the fault type is a hardware fault, the preset threshold is adjusted to 0.7; if the fault type is a software configuration fault, the preset threshold is adjusted to 0.8.
[0073] This embodiment adopts a dynamic threshold adjustment mechanism, which adjusts the value of the preset threshold according to the fault type, thereby reducing the problem of missed selection and misselection caused by fixed thresholds and improving the accuracy of candidate root causes.
[0074] In some embodiments, determining the similarity between the target fault and the historical fault based on each of the matching degrees includes:
[0075] The first target value is obtained by weighted summation of the matching degrees described above.
[0076] The penalty term is determined based on the penalty coefficient and the cosine similarity of each term.
[0077] A second target value is determined based on the first target value and the penalty term;
[0078] The ratio of the first target value to the second target value is calculated to obtain the similarity between the target fault and the historical fault.
[0079] Based on the feature importance of each feature attribute, the matching degrees are weighted and summed, and the resulting weighted sum is called the first target value. In an embodiment where the fault's feature attributes include fault type, fault qualifiers, fault trigger, fault source, fault impact range, fault duration, fault detectability, and fault-related resources, the weighted sum of the matching degrees based on the feature importance of each feature attribute can be expressed as:
[0080] .
[0081] Let represent the feature importance corresponding to feature attribute j. Initially, feature importance is evenly distributed, with each feature attribute having an importance of 1 / 8. Special importance for feature attributes can be adjusted. For example, if fault qualifiers contain a lot of noise, the feature importance corresponding to fault qualifiers can be reduced.
[0082] The method for determining the penalty term based on the penalty coefficient and each cosine similarity can be as follows:
[0083] Calculate the difference between the target value and the cosine similarity;
[0084] Based on the feature importance corresponding to the feature attribute, the squares of each difference are weighted and summed.
[0085] Multiplying the weighted sum by the penalty coefficient yields the penalty term. The target value can be 1.
[0086] The penalty coefficient defaults to 1 and is adjustable. The penalty item is represented as follows: λ is the penalty coefficient. (Through...) It can strengthen the penalty for low similarity feature attributes, while combining Weighting can make the penalty for mismatch of important feature attributes stronger.
[0087] The second target value can be determined by taking the square root of the sum of the square of the first target value and the square of the penalty term.
[0088] Calculate the ratio of the first target value to the second target value to obtain the similarity between the target fault and the historical fault. Similarity between target fault X and historical fault Y The calculation method is as follows:
[0089] .
[0090] S102: Select the number of historical faults with the highest similarity to obtain a set of similar faults.
[0091] After executing step S101 and determining the similarity between the target fault and all historical faults in the fault database, the similarity scores are sorted in descending order, and the top k historical faults are selected. The selected top k historical faults form a set of similar faults.
[0092] S103: Construct a fault tree based on the set of similar faults; the fault phenomena of historical faults in the set of similar faults are used as the top events of the fault tree, and the root causes of historical faults in the set of similar faults are used as the basic events of the fault tree.
[0093] This embodiment uses operating system-level fault phenomena (such as operating system crashes, operating system boot failures, or critical system service interruptions) as the top events in the fault tree. During fault phenomenon analysis and troubleshooting, specific hardware failures, software defects, or configuration errors are located—that is, the lowest-level, indivisible fault causes are identified. These indivisible fault causes, or root causes, are the basic events in the fault tree. For each historical fault in a similar fault set, a corresponding fault tree is constructed.
[0094] In some embodiments, constructing a fault tree based on the set of similar faults includes:
[0095] The fault tree is constructed by using the fault phenomena of historical faults in the similar fault set as the top events and the root causes of historical faults in the similar fault set as the basic events. The logic gates are used to describe the causal relationships between events.
[0096] For each historical fault in a similar fault set, a fault tree is constructed using its fault phenomenon (e.g., "operating system crash") as the top event and its indivisible root cause (e.g., "configuration file syntax error" or "hardware damage") as the base events, through logic gates (AND gates / OR gates). Logic gates are used to describe the causal relationships between events (e.g., an "AND gate" indicates that multiple events must occur simultaneously to trigger a higher-level event).
[0097] Fault tree templates can be reused directly, reducing the cost of repetitive modeling.
[0098] S104: Extract the minimum cut set of the fault tree; the minimum cut set includes several of the basic events.
[0099] A cut set is a set of basic events that, if all of these basic events occur simultaneously, will cause the top event to occur. A minimal cut set is the minimum set of basic events that trigger the top event. If any basic event in the minimal cut set is removed, the top event will not occur.
[0100] In some embodiments, extracting the minimum cut set of the fault tree includes:
[0101] Step 1: Take the i-th cut set and determine whether i is less than the number of cut sets;
[0102] Step 2: If i equals the number of cut sets, then the i-th cut set is the minimum cut set;
[0103] Step 3: If i is less than the number of cut sets, then take the j-th cut set and determine whether the i-th cut set contains the j-th cut set;
[0104] Step 4: If the i-th cut set does not contain the j-th cut set, then update j to j plus one, and if the updated j is less than the number of cut sets, return to step 3;
[0105] Step 5: If the i-th cut set contains the j-th cut set, delete the i-th cut set, decrease the number of cut sets by one, update i to i plus 1, and then execute step 1.
[0106] This embodiment uses a row-column method to extract the minimum cut set of the fault tree. Starting from the top event, each "OR gate" is replaced vertically (making the sub-event an independent cut set element), and each "AND gate" is replaced horizontally (combining the sub-events into a cut set element). The minimum cut set is obtained by Boolean algebra simplification.
[0107] refer to Figure 2 As shown, remove duplicate nodes from each cut set, assign the value i=0, and assign a count value equal to the number of cut sets. Take the i-th cut set and check if i is less than the count value. If i is less than the count value, assign j=0 and take the j-th cut set. Check if the i-th cut set contains the j-th cut set. If the i-th cut set contains the j-th cut set, delete the i-th cut set, decrement the count value, increment i, and return to the step of checking if i is less than the count value. If the i-th cut set does not contain the j-th cut set, increment j and check if j is less than the count value. If j is less than the count value, return to the step of checking if j-th cut set. If j is not less than the count value, increment i and return to the step of checking if i is less than the count value. Continue until i is not less than the count value; at this point, the i-th cut set is the minimum cut set.
[0108] S105: Determine the importance of the basic events in the minimum cut set, and select the basic event with the highest importance as the candidate root cause of the target fault.
[0109] Candidate root causes refer to the fundamental reasons that may cause the target failure. The various candidate root causes of the target failure constitute the root cause candidate set, which is the set of the most likely root causes of the target failure. Users can use this candidate set to refine the localization scope and refer to historical solutions in the fault knowledge base to resolve the fault.
[0110] In some embodiments, determining the importance of basic events in the minimal cut set includes:
[0111] The structural importance, time-sensitivity importance, and propagation risk of the basic events in the minimum cut set are determined. The structural importance characterizes the degree of influence of the basic events on the occurrence of the top event. The time-sensitivity importance characterizes the impact of the time cost of event repair on system availability. The propagation risk characterizes the cascading risk of events in the system.
[0112] The importance of the basic events in the minimum cut set is obtained by weighted summation of the structural importance, timeliness importance, and propagation risk.
[0113] This embodiment combines structural importance, timeliness importance, and propagation risk to determine the importance of basic events in the minimum cut set, comprehensively and accurately assessing the importance of basic events and reducing the root cause locking error rate.
[0114] The importance of basic events is calculated as follows:
[0115] .
[0116] Where x is a basic event, and the coefficients α, β, γ satisfy α + β + γ = 1. This represents the structural importance of the basic event x. The numerical value of the structural importance reflects the degree of influence of the basic event x on the occurrence of the top event. The higher the value, the greater the influence of the occurrence of the basic event x on the occurrence of the top event. Let x be the timeliness importance of the basic event. Timeliness importance measures the impact of the time cost of event repair on system availability. Generally, the shorter the repair time, the higher the importance (i.e., the higher the urgency). To determine the propagation risk of a basic event x, a random walk algorithm can be used based on a fault propagation graph model to calculate the propagation risk of the basic event x, capturing cascading effects at different levels (such as the propagation phenomenon where a hardware failure like bad sectors on a hard drive leads to file system corruption, which in turn leads to service crashes).
[0117] refer to Figure 3As shown, the target fault is input, and the similarity between the target fault and historical faults is calculated. Based on the similarity, the top k historical faults are selected. A fault tree is constructed based on the selected historical faults. Minimum cut sets are extracted, and the importance of basic events is determined based on their structural importance, timeliness importance, and propagation risk. Based on the importance of each basic event, the top k possible causes of the fault are determined.
[0118] In some embodiments, determining the structural importance of basic events in the minimum cut set includes:
[0119] The structural importance of a basic event is determined based on the number of minimal cut sets containing the basic event and the number of basic events in the minimal cut sets.
[0120] The structural importance of a basic event x can be calculated as follows:
[0121] .
[0122] in, Let x be the minimum cut set containing the basic event x. for The number of basic events.
[0123] For example, the minimal cut set containing the basic event x is and , The number of basic events is 5. If the number of basic events is 7, then .
[0124] In some embodiments, determining the timeliness importance of basic events in the minimal cut set includes:
[0125] The timeliness importance of the basic events is determined based on the average repair time of the basic events and a preset smoothing factor.
[0126] The timeliness importance of a basic event x can be calculated as follows:
[0127] .
[0128] Where MTTR(x) is the average recovery time of basic event x. This is a smoothing factor used to prevent boundary overflow.
[0129] In some embodiments, it also includes:
[0130] Cluster analysis is performed on each of the candidate root causes to merge duplicate or strongly associated candidate root causes.
[0131] Cluster analysis of each candidate root cause, merging duplicate or strongly correlated root causes (such as "configuration file error" and "syntax error"), can further refine the scope of the problem.
[0132] In summary, this invention determines the similarity between the target fault and historical faults based on multiple dimensions of feature attributes, avoiding the one-sidedness of similarity calculation based on a single dimension, improving the accuracy of similarity calculation, and thus improving the accuracy of similarity-based fault screening. Furthermore, this invention constructs a fault tree corresponding to historical faults and extracts the minimum cut set to determine the candidate root causes of the target fault, thereby avoiding redundant investigation and improving fault location efficiency.
[0133] The fault location method provided in the above embodiments of the present invention can be applied not only to operating system fault location, but also to fields such as industrial equipment diagnosis and network fault diagnosis. In the field of industrial equipment diagnosis, the appropriate fault source refers to replacing the faulty module with the faulty equipment component, and the fault triggering condition with operating parameter thresholds. Operating parameter thresholds are used for fault location in assembly line machinery. In the field of network fault diagnosis, the fault type can be extended to the network layer, such as the physical layer or protocol layer. Furthermore, network fault location can be achieved by combining routing information.
[0134] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method.
[0135] Embodiments of the present invention also provide a fault location device, with reference to Figure 4 As shown, the device includes:
[0136] The similarity determination module 10 is used to determine the similarity between the target fault and the historical fault based on the feature attributes of the target fault in multiple dimensions and the feature attributes of the historical fault in multiple dimensions.
[0137] The historical fault selection module 20 is used to select the top preset number of historical faults with the highest similarity to obtain a set of similar faults;
[0138] The fault tree construction module 30 is used to construct a fault tree based on the set of similar faults; the fault phenomena of historical faults in the set of similar faults are used as the top events of the fault tree, and the root causes of historical faults in the set of similar faults are used as the basic events of the fault tree.
[0139] Minimum cut set extraction module 40 is used to extract the minimum cut set of the fault tree; the minimum cut set includes several of the basic events;
[0140] The root cause determination module 50 is used to determine the importance of basic events in the minimum cut set and select the basic event with the highest importance as the candidate root cause of the target fault.
[0141] Based on the above embodiments, as a specific implementation method, the similarity determination module 10 is used for:
[0142] Based on the fault type, fault qualifiers, fault trigger, fault source, fault impact range, fault duration, fault detectability, and fault-related resources of the target fault, and the fault type, fault qualifiers, fault trigger, fault source, fault impact range, fault duration, fault detectability, and fault-related resources of the historical fault, the similarity between the target fault and the historical fault is determined.
[0143] Based on the above embodiments, as a specific implementation method, the root cause determination module 50 includes:
[0144] A determining unit is used to determine the structural importance, timeliness importance, and propagation risk of basic events in the minimum cut set; the structural importance characterizes the degree of influence of basic events on the occurrence of the top event; the timeliness importance characterizes the impact of the time cost of event repair on system availability; and the propagation risk characterizes the cascading risk of events in the system.
[0145] The calculation unit is used to perform a weighted summation of the structural importance, timeliness importance, and propagation risk to obtain the importance of the basic events in the minimum cut set.
[0146] Based on the above embodiments, as a specific implementation method, the similarity determination module 10 includes:
[0147] A matching degree determination unit is used to determine the matching degree between the feature attributes of the target fault in each dimension and the feature attributes corresponding to the historical fault.
[0148] A similarity determination unit is used to determine the similarity between the target fault and the historical fault based on each of the matching degrees.
[0149] Based on the above embodiments, as a specific implementation method, the matching degree determination unit includes:
[0150] The cosine similarity determination unit is used to determine the cosine similarity between the feature attributes of the target fault and the feature attributes corresponding to the historical faults.
[0151] The comparison unit is used to compare the cosine similarity with a preset threshold.
[0152] The first determining unit is configured to determine that if the cosine similarity is greater than or equal to the preset threshold, the matching degree between the feature attributes of the target fault and the feature attributes corresponding to the historical fault is equal to the cosine similarity.
[0153] The second determining unit is configured to determine that if the cosine similarity is less than the preset threshold, the matching degree between the feature attributes of the target fault and the feature attributes corresponding to the historical fault is equal to zero.
[0154] Based on the above embodiments, as a specific implementation method, the similarity determination unit includes:
[0155] The first calculation unit is used to perform a weighted summation of the matching degrees to obtain a first target value;
[0156] The second calculation unit is used to determine the penalty term based on the penalty coefficient and each of the cosine similarities;
[0157] The third calculation unit is used to determine the second target value based on the first target value and the penalty term;
[0158] The fourth calculation unit is used to calculate the ratio of the first target value to the second target value to obtain the similarity between the target fault and the historical fault.
[0159] Based on the above embodiments, as a specific implementation method, the fault tree construction module 30 is used for:
[0160] The fault tree is constructed by using the fault phenomena of historical faults in the similar fault set as the top events and the root causes of historical faults in the similar fault set as the basic events. The logic gates are used to describe the causal relationships between events.
[0161] Based on the above embodiments, as a specific implementation method, the minimum cut set extraction module 40 is used to perform the following steps:
[0162] Step 1: Take the i-th cut set and determine whether i is less than the number of cut sets;
[0163] Step 2: If i equals the number of cut sets, then the i-th cut set is the minimum cut set;
[0164] Step 3: If i is less than the number of cut sets, then take the j-th cut set and determine whether the i-th cut set contains the j-th cut set;
[0165] Step 4: If the i-th cut set does not contain the j-th cut set, then update j to j plus one, and if the updated j is less than the number of cut sets, return to step 3;
[0166] Step 5: If the i-th cut set contains the j-th cut set, delete the i-th cut set, decrease the number of cut sets by one, update i to i plus 1, and then execute step 1.
[0167] Based on the above embodiments, as a specific implementation method, the determining unit is used for:
[0168] The structural importance of a basic event is determined based on the number of minimal cut sets containing the basic event and the number of basic events in the minimal cut sets.
[0169] Based on the above embodiments, as a specific implementation method, the determining unit is used for:
[0170] The timeliness importance of the basic events is determined based on the average repair time of the basic events and a preset smoothing factor.
[0171] Based on the above embodiments, as a specific implementation method, it further includes:
[0172] The threshold adjustment module is used to adjust the value of the preset threshold according to the fault type of the target fault.
[0173] Based on the above embodiments, as a specific implementation method, it further includes:
[0174] The merging module is used to perform cluster analysis on each of the candidate root causes, and merge duplicate candidate root causes or strongly associated candidate root causes.
[0175] For a description of the features in the embodiment corresponding to the fault location device, please refer to the relevant description of the embodiment corresponding to the fault location method, which will not be repeated here.
[0176] Embodiments of the present invention also provide an electronic device, with reference to Figure 5 As shown, the electronic device includes a memory 1 and a processor 2. The memory 1 stores a computer program, and the processor 2 is configured to run the computer program to perform the steps in any of the above-described fault location method embodiments.
[0177] Embodiments of the present invention also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described fault location method embodiments when it is run.
[0178] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0179] Embodiments of the present invention also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described fault location method embodiments.
[0180] Embodiments of the present invention also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described fault location method embodiments.
[0181] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0182] The foregoing has provided a detailed description of the fault location method, electronic device, storage medium, and program product provided by this invention. Specific examples have been used to illustrate the principles and implementation methods of this invention. The descriptions of the above embodiments are only intended to help understand the method and core ideas of this invention. It should be noted that those skilled in the art can make various improvements and modifications to this invention without departing from its principles, and these improvements and modifications also fall within the protection scope of this invention.
Claims
1. A fault location method, characterized in that, include: The similarity between the target fault and the historical fault is determined based on the multi-dimensional feature attributes of the target fault and the multi-dimensional feature attributes of the historical fault. Select the first preset number of historical faults with the highest similarity to obtain a set of similar faults; Construct a fault tree based on the set of similar faults; The fault phenomena of historical faults in the similar fault set are taken as the top events of the fault tree, and the root causes of historical faults in the similar fault set are taken as the basic events of the fault tree. Extract the minimum cut set of the fault tree; the minimum cut set includes several of the basic events; Determine the importance of the basic events in the minimum cut set, and select the basic event with the highest importance as the candidate root cause of the target fault; Determining the importance of basic events in the minimum cut set includes: The structural importance, time-sensitivity importance, and propagation risk of the basic events in the minimum cut set are determined. The structural importance characterizes the degree of influence of the basic events on the occurrence of the top event. The time-sensitivity importance characterizes the impact of the time cost of event repair on system availability. The propagation risk characterizes the cascading risk of events in the system. The importance of the basic events in the minimum cut set is obtained by weighted summation of the structural importance, timeliness importance, and propagation risk. Determining the structural importance of basic events in the minimum cut set includes: The structural importance of the basic events is determined based on the number of minimal cut sets containing the basic events and the number of basic events in the minimal cut sets. Determining the timeliness importance of basic events in the minimum cut set includes: The timeliness importance of the basic events is determined based on the average repair time of the basic events and a preset smoothing factor.
2. The fault location method according to claim 1, characterized in that, Determining the similarity between the target fault and the historical faults based on multiple dimensions of feature attributes of the target fault and multiple dimensions of feature attributes of historical faults includes: Determine the matching degree between the feature attributes of the target fault in each dimension and the feature attributes corresponding to the historical fault; Based on the matching degree, the similarity between the target fault and the historical fault is determined.
3. The fault location method according to claim 2, characterized in that, Determining the matching degree between the feature attributes of the target fault in each dimension and the feature attributes corresponding to the historical faults includes: Determine the cosine similarity between the feature attributes of the target fault and the feature attributes corresponding to the historical faults; Compare the cosine similarity with a preset threshold; If the cosine similarity is greater than or equal to the preset threshold, then the matching degree between the feature attributes of the target fault and the feature attributes corresponding to the historical fault is equal to the cosine similarity. If the cosine similarity is less than the preset threshold, then the matching degree between the feature attributes of the target fault and the feature attributes corresponding to the historical fault is zero.
4. The fault location method according to claim 3, characterized in that, Determining the similarity between the target fault and the historical fault based on each of the aforementioned matching degrees includes: The first target value is obtained by weighted summation of the matching degrees described above. The penalty term is determined based on the penalty coefficient and the cosine similarity of each term. A second target value is determined based on the first target value and the penalty term; The ratio of the first target value to the second target value is calculated to obtain the similarity between the target fault and the historical fault.
5. The fault location method according to claim 1, characterized in that, Constructing a fault tree based on the set of similar faults includes: The fault tree is constructed by using the fault phenomena of historical faults in the similar fault set as the top events and the root causes of historical faults in the similar fault set as the basic events. The logic gates are used to describe the causal relationships between events.
6. The fault location method according to claim 1, characterized in that, Extracting the minimum cut set of the fault tree includes: Step 1: Take the i-th cut set and determine whether i is less than the number of cut sets; Step 2: If i equals the number of cut sets, then the i-th cut set is the minimum cut set; Step 3: If i is less than the number of cut sets, then take the j-th cut set and determine whether the i-th cut set contains the j-th cut set; Step 4: If the i-th cut set does not contain the j-th cut set, then update j to j plus one, and if the updated j is less than the number of cut sets, return to step 3; Step 5: If the i-th cut set contains the j-th cut set, delete the i-th cut set, decrease the number of cut sets by one, update i to i plus 1, and then execute step 1.
7. The fault location method according to claim 3, characterized in that, Also includes: Adjust the value of the preset threshold according to the fault type of the target fault.
8. The fault location method according to claim 1, characterized in that, Determining the similarity between the target fault and the historical faults based on multiple dimensions of feature attributes of the target fault and multiple dimensions of feature attributes of historical faults includes: Based on the fault type, fault qualifiers, fault trigger, fault source, fault impact range, fault duration, fault detectability, and fault-related resources of the target fault, and the fault type, fault qualifiers, fault trigger, fault source, fault impact range, fault duration, fault detectability, and fault-related resources of the historical fault, the similarity between the target fault and the historical fault is determined.
9. The fault location method according to claim 1, characterized in that, Also includes: Cluster analysis was performed on each of the candidate root causes to merge duplicate candidate root causes.
10. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the fault location method as described in any one of claims 1 to 9 when executing the computer program.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, it implements the steps of the fault location method as described in any one of claims 1 to 9.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the fault location method as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Remote monitoring method for operating state of numerical control machine tool
CN108196514A
Helicopter intelligent fault diagnosis method based on fault tree and case-based reasoning
CN111581739A