A fault handling method and apparatus

By identifying and matching abnormal indicators of business systems, and using loss mitigation measures for known fault scenarios or training models to handle unknown fault scenarios, the problem of low fault handling efficiency in existing technologies is solved, and efficient fault recovery and system operation assurance are achieved.

CN114138610BActive Publication Date: 2025-11-18CHINA CONSTRUCTION BANK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111485663.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-07
Publication Date
2025-11-18
Estimated Expiration
2041-12-07

AI Technical Summary

Technical Problem

Existing technologies are unable to effectively handle faults in business systems, leading to production and transaction service interruptions and impacting business processing efficiency.

Method used

By responding to fault diagnosis commands, abnormal indicators of the target system can be obtained, key indicators can be identified, known fault scenarios can be matched and loss mitigation measures can be taken, or a trained scenario matching model can be used to handle unknown fault scenarios.

Benefits of technology

It improved fault handling efficiency, ensured the normal operation of the system and business processing efficiency, reduced manual workload, and improved the accuracy and efficiency of fault identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114138610B_ABST
    Figure CN114138610B_ABST
Patent Text Reader

Abstract

The application discloses a fault processing method and device, which can obtain at least one abnormal index currently appeared in a target system in response to a fault diagnosis instruction, identify at least one key index from the abnormal indexes, determine whether a first fault scene of the target system currently matches at least one known fault scene in a known fault scene group based on each identified key index, determine a first known fault scene matched with the first fault scene from the known fault scene group if yes, and perform fault processing by using a stop-loss processing mode corresponding to the first known fault scene, so that the fault processing efficiency can be effectively improved, and the system operation efficiency can be ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of fault handling, and more particularly to a fault handling method and apparatus. Background Technology

[0002] As infrastructure scales up, the architectural complexity of business systems continues to increase.

[0003] If a business system experiences a failure that affects production, transactions, or causes service interruption during operation, the failure needs to be addressed as soon as possible to restore normal operation and ensure business processing efficiency.

[0004] However, existing technologies are unable to effectively handle faults. Summary of the Invention

[0005] In view of the above problems, the present invention provides a fault handling method and apparatus to overcome or at least partially solve the above problems, the technical solution of which is as follows:

[0006] A fault handling method, comprising:

[0007] In response to a fault diagnosis command, obtain at least one abnormal indicator that has appeared in the target system.

[0008] Identify at least one key indicator from each of the aforementioned abnormal indicators;

[0009] Based on each of the identified key indicators, determine whether the current first fault scenario of the target system matches at least one known fault scenario in the known fault scenario group;

[0010] If so, a first known fault scenario matching the first fault scenario is determined from the known fault scenario group, and the fault is handled using the loss prevention handling method corresponding to the first known fault scenario.

[0011] Optionally, the method further includes:

[0012] If the first fault scenario does not match any known fault scenario in the known fault scenario group, then the first fault scenario is determined to be an unknown fault scenario.

[0013] Using a trained scene matching model, a second known fault scene that matches the first fault scene is determined from the known fault scene group, and the fault is handled using the loss prevention handling method corresponding to the second known fault scene.

[0014] Optionally, based on a first key indicator, determining whether the first fault scenario matches at least one known fault scenario in the known fault scenario group includes:

[0015] Determine the third known fault scenario corresponding to the first key indicator; wherein, the third known fault scenario includes at least one predefined scenario indicator;

[0016] Determine whether any of the aforementioned abnormal indicators are other than the first key indicator;

[0017] If so, then among the abnormal indicators, determine the occurrence time of each of the scenario indicators other than the first key indicator, and determine the earliest occurrence time from among the occurrence times;

[0018] The difference obtained by subtracting the earliest occurrence time from the occurrence time of the first key indicator is determined as the first difference;

[0019] Based on the first difference, it is determined whether the first fault scenario matches the third known fault scenario.

[0020] Optionally, determining whether the first fault scenario matches the third known fault scenario based on the first difference includes:

[0021] If the first difference is greater than a predefined duration, then it is determined that the first fault scenario does not match the third known fault scenario.

[0022] Optionally, determining whether the first fault scenario matches the third known fault scenario based on the first difference includes:

[0023] If the first difference is not less than zero and not greater than the predefined duration, determine whether each of the abnormal indicators has included all the scenario indicators in the third known fault scenario;

[0024] If so, then the first fault scenario is determined to match the third known fault scenario;

[0025] Otherwise, within the predefined duration starting from the earliest occurrence time, it is detected in the target system whether all the scenario indicators in the third known fault scenario appear. If so, it is determined that the first fault scenario matches the third known fault scenario; otherwise, it is determined that the first fault scenario does not match the third known fault scenario.

[0026] Optionally, determining whether the first fault scenario matches the third known fault scenario based on the first difference includes:

[0027] If the first difference is less than zero, then within the predefined time period starting from the time the first key indicator appears, it is detected in the target system whether all the scenario indicators in the third known fault scenario appear. If so, it is determined that the first fault scenario matches the third known fault scenario; otherwise, it is determined that the first fault scenario does not match the third known fault scenario.

[0028] Optionally, the method further includes:

[0029] If none of the scenario indicators other than the first key indicator exist among the abnormal indicators, then within the predefined time period starting from the occurrence time of the first key indicator, it is detected in the target system whether all the scenario indicators in the third known fault scenario appear. If so, it is determined that the first fault scenario matches the third known fault scenario; otherwise, it is determined that the first fault scenario does not match the third known fault scenario.

[0030] A fault handling apparatus includes: a first obtaining unit, a first identifying unit, a first determining unit, a second determining unit, and a first processing unit; wherein:

[0031] The first obtaining unit is configured to obtain at least one abnormal indicator that has appeared in the target system in response to a fault diagnosis command;

[0032] The first identification unit is used to identify at least one key indicator from each of the abnormal indicators;

[0033] The first determining unit is used to determine, based on each of the identified key indicators, whether the current first fault scenario of the target system matches at least one known fault scenario in the known fault scenario group; if so, the second determining unit is triggered.

[0034] The second determining unit is used to determine a first known fault scenario that matches the first fault scenario from the known fault scenario group;

[0035] The first processing unit is used to perform fault handling using a loss mitigation method corresponding to the first known fault scenario.

[0036] Optionally, the device further includes: a third determining unit, a fourth determining unit, and a second processing unit;

[0037] The third determining unit is configured to determine the first fault scenario as an unknown fault scenario if the first fault scenario does not match any known fault scenario in the known fault scenario group.

[0038] The fourth determining unit is used to use a trained scene matching model to determine a second known fault scene that matches the first fault scene from the known fault scene group.

[0039] The second processing unit is used to perform fault handling using a loss mitigation method corresponding to the second known fault scenario.

[0040] Optionally, based on a first key indicator, determining whether the first fault scenario matches at least one known fault scenario in the known fault scenario group is configured as follows:

[0041] Determine the third known fault scenario corresponding to the first key indicator; wherein, the third known fault scenario includes at least one predefined scenario indicator;

[0042] Determine whether any of the aforementioned abnormal indicators are other than the first key indicator;

[0043] If so, then among the abnormal indicators, determine the occurrence time of each of the scenario indicators other than the first key indicator, and determine the earliest occurrence time from among the occurrence times;

[0044] The difference obtained by subtracting the earliest occurrence time from the occurrence time of the first key indicator is determined as the first difference;

[0045] Based on the first difference, it is determined whether the first fault scenario matches the third known fault scenario.

[0046] Optionally, determining whether the first fault scenario matches the third known fault scenario based on the first difference is configured as follows:

[0047] If the first difference is greater than a predefined duration, then it is determined that the first fault scenario does not match the third known fault scenario.

[0048] Optionally, determining whether the first fault scenario matches the third known fault scenario based on the first difference is configured as follows:

[0049] If the first difference is not less than zero and not greater than the predefined duration, determine whether each of the abnormal indicators has included all the scenario indicators in the third known fault scenario;

[0050] If so, then the first fault scenario is determined to match the third known fault scenario;

[0051] Otherwise, within the predefined duration starting from the earliest occurrence time, it is detected in the target system whether all the scenario indicators in the third known fault scenario appear. If so, it is determined that the first fault scenario matches the third known fault scenario; otherwise, it is determined that the first fault scenario does not match the third known fault scenario.

[0052] Optionally, determining whether the first fault scenario matches the third known fault scenario based on the first difference is configured as follows:

[0053] If the first difference is less than zero, then within the predefined time period starting from the time the first key indicator appears, it is detected in the target system whether all the scenario indicators in the third known fault scenario appear. If so, it is determined that the first fault scenario matches the third known fault scenario; otherwise, it is determined that the first fault scenario does not match the third known fault scenario.

[0054] Optionally, the device further includes: a detection unit, a fifth determining unit, and a sixth determining unit; wherein:

[0055] The detection unit is configured to, if none of the scenario indicators other than the first key indicator exist among the abnormal indicators, detect whether all the scenario indicators in the third known fault scenario appear in the target system within the predefined time period starting from the time the first key indicator appears; if so, trigger the fifth determination unit; otherwise, trigger the sixth determination unit.

[0056] The fifth determining unit is used to determine whether the first fault scenario matches the third known fault scenario;

[0057] The sixth determining unit is used to determine that the first fault scenario does not match the third known fault scenario.

[0058] The fault handling method and apparatus proposed in this invention can respond to fault diagnosis instructions, obtain at least one abnormal indicator that has appeared in the target system, identify at least one key indicator from each abnormal indicator, and determine whether the current first fault scenario of the target system matches at least one known fault scenario in the known fault scenario group based on each identified key indicator. If so, the first known fault scenario that matches the first fault scenario is determined from the known fault scenario group, and the fault is handled using the loss mitigation handling method corresponding to the first known fault scenario. This can effectively improve fault handling efficiency and ensure system operating efficiency.

[0059] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention, it can be implemented according to the contents of the specification. In order to make the above and other objects, features and advantages of the present invention more obvious and understandable, specific embodiments of the present invention are described below. Attached Figure Description

[0060] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0061] Figure 1 A flowchart of a first fault handling method provided by an embodiment of the present invention is shown;

[0062] Figure 2 A flowchart of a second fault handling method provided in an embodiment of the present invention is shown;

[0063] Figure 3 A schematic diagram of a barrier-clearing tree structure provided by an embodiment of the present invention is shown;

[0064] Figure 4 A schematic diagram of the structure of the first fault handling device provided in an embodiment of the present invention is shown. Detailed Implementation

[0065] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this invention will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art.

[0066] like Figure 1 As shown, this embodiment proposes a first fault handling method, which may include the following steps:

[0067] S101. In response to the fault diagnosis command, obtain at least one abnormal indicator that has appeared in the target system.

[0068] It should be noted that this invention can be applied to electronic devices, such as desktop computers, tablet computers, and servers.

[0069] The fault diagnosis command can be an instruction that triggers the aforementioned electronic device to perform fault diagnosis. It should be noted that the fault diagnosis command can be generated by the aforementioned electronic device when it detects abnormal indicators, or it can be input by a fault monitoring system, other devices, or manually; this invention does not limit this. The fault monitoring system can be installed in the aforementioned electronic device or in other devices.

[0070] The target system can be a business system for which fault handling is required in this invention, such as a banking business system and a ticket sales system.

[0071] Among them, abnormal indicators can be abnormal operating indicators of a certain operation and maintenance object in the target system.

[0072] In this context, the operation and maintenance objects can be the objects that need to be managed and maintained during the operation of the target system. For example, when the target system is a banking system, the operation and maintenance objects could be Oracle databases and Weblogic middleware, etc.

[0073] Optionally, the present invention can generate a fault diagnosis command when a specified type of abnormal indicator is detected in the target system during the operation of the target system, and then, in response to the fault diagnosis command, determine all the abnormal indicators that have appeared in the target system.

[0074] Optionally, the present invention can obtain all abnormal indicators that have appeared in the target system when receiving a fault diagnosis instruction sent by a fault monitoring system or other devices.

[0075] S102. Identify at least one key indicator from among the abnormal indicators;

[0076] Among them, key indicators can be important operational metrics for measuring whether a certain failure scenario has occurred in the target system.

[0077] It should be noted that key performance indicators (KPIs) can be operational indicators that are pre-specified by technical personnel.

[0078] S103. Based on the identified key indicators, determine whether the current first fault scenario of the target system matches at least one known fault scenario in the known fault scenario group. If so, proceed to step S104.

[0079] The first fault scenario can be the fault scenario currently occurring in the target system that needs to be identified. It should be noted that the present invention can respond to a fault diagnosis command and begin to diagnose and identify the first fault scenario even before the fault phenomena and abnormal indicators of the first fault scenario are fully manifested.

[0080] Specifically, a fault scenario can be identified by an anomalous indicator set that includes one or more abnormal indicators. When all the abnormal indicators in a certain anomalous indicator set that can identify a certain fault scenario appear in the target system, the present invention can determine that the target system is currently experiencing that fault scenario.

[0081] Among them, the known fault scenario group can be a set consisting of multiple known fault scenarios.

[0082] Among them, known fault scenarios can be fault scenarios that have occurred in the past, whose entire fault phenomenon has been recorded, and whose fault handling methods have been verified.

[0083] It is understandable that the set of abnormal indicators corresponding to a known fault scenario can be known, that is, all abnormal indicators in the set of abnormal indicators corresponding to a known fault scenario can be known.

[0084] Specifically, when determining whether a first fault scenario matches a known fault scenario, the present invention can determine whether the first fault scenario of the target system is a known fault scenario corresponding to the abnormal indicator set by determining whether the target system exhibits all abnormal indicators in the abnormal indicator set corresponding to the known fault scenario within a specified time period.

[0085] S104. Determine the first known fault scenario that matches the first fault scenario from the known fault scenario group;

[0086] The first known fault scenario can be a known fault scenario that matches the first fault scenario in the known fault scenario group.

[0087] Specifically, if the first fault scenario matches at least one known fault scenario in the known fault scenario group, the present invention can determine the first known fault scenario that matches the first fault scenario in the known fault scenario group.

[0088] It should be noted that if the first fault scenario matches multiple known fault scenarios in the known fault scenario group, the present invention can randomly select one of the multiple matching known fault scenarios as the first known fault scenario, or it can be manually selected by a technician.

[0089] S105. Use the loss-stopping method corresponding to the first known fault scenario to handle the fault.

[0090] Among them, the loss mitigation measures can be used to handle and recover from the failures that occur in the target system, such as service restart, fault isolation, service switching, version rollback, backup and recovery, parameter modification, flow control, cluster expansion, service circuit breaking and / or service degradation.

[0091] Specifically, this invention can restore the system to its faults by using a loss mitigation method, thereby resolving and restoring the business, transaction and service interruptions suffered by the target system due to the fault scenario to a certain extent.

[0092] It is understandable that the loss mitigation methods for known fault scenarios can be known. Therefore, when the present invention determines that the first fault scenario matches the first known fault scenario, it can utilize the loss mitigation methods corresponding to the first known fault scenario to handle the fault in the first fault scenario. This allows the target system to effectively resolve and recover from the fault caused by the first fault scenario, thereby effectively improving fault handling efficiency and ensuring system operating efficiency.

[0093] It should also be noted that traditional fault handling models rely on documented emergency plans. This approach presents several problems. First, documented emergency plans inevitably lead to inefficiency. In production emergencies, it's not uncommon to manually search for solutions in emergency plans offline even when the fault scenario is clear. This process may involve multiple communications between technical personnel, resulting in significant uncertainty in fault handling time. If the system administrator (A role) is not on-site, there are even more uncontrollable factors, further reducing efficiency. Second, the quality of handling is heavily dependent on the quality of the emergency plan, making it highly reliant on the system administrator who wrote it. System administrators with strong technical skills and extensive experience are likely to write high-quality emergency plans, while those with little experience in incident handling may be less proficient in fault location and recovery, affecting the quality of handling. This invention, however, addresses this issue by... Figure 1 The method shown can automatically identify and handle fault scenarios, effectively reducing manual workload, ensuring the accuracy and efficiency of fault scenario identification, and improving fault handling efficiency.

[0094] The fault handling method proposed in this embodiment can respond to fault diagnosis instructions, obtain at least one abnormal indicator that has appeared in the target system, identify at least one key indicator from each abnormal indicator, and determine whether the current first fault scenario of the target system matches at least one known fault scenario in the known fault scenario group based on each identified key indicator. If so, the first known fault scenario that matches the first fault scenario is determined from the known fault scenario group, and the fault is handled using the loss mitigation handling method corresponding to the first known fault scenario. This can effectively improve fault handling efficiency and ensure system operating efficiency.

[0095] based on Figure 1 As shown, Figure 2As shown, this embodiment proposes a second fault handling method. This method may further include the following steps:

[0096] S201. If the first fault scenario does not match any known fault scenario in the known fault scenario group, then the first fault scenario is determined to be an unknown fault scenario.

[0097] Among them, unknown fault scenarios can be fault scenarios that are not known fault scenarios.

[0098] Specifically, the present invention can determine the first fault scenario as an unknown fault scenario when none of the known fault scenarios in the first fault scenario unknown fault scenario group match.

[0099] S202. Using the trained scene matching model, determine the second known fault scene that matches the first fault scene from the known fault scene group, and use the loss prevention and control method corresponding to the second known fault scene to handle the fault.

[0100] It should be noted that this invention treats the identification of unknown fault scenarios as a text classification problem. It vectorizes the event phenomena of known fault scenarios, which consist of elements such as maintenance objects, abnormal indicators, normal indicators, and indicator values. Then, it defines a text similarity metric function and trains a scene matching model to classify unknown fault scenarios, identifying known fault scenarios with a high degree of similarity to the fault phenomena or features of the unknown fault scenarios. It is understood that the scene matching model can be a text classification model, such as a K-Nearest Neighbor (KNN) classification model or a support vector machine (SVM) classification model.

[0101] A good metric function can increase intra-class similarity and decrease inter-class similarity. In this embodiment, the scene matching model needs to identify unknown fault scenarios based on both semantics and the actual similarity of the faults. For example, Weblogic FullGC and OutOfMemoryError often occur together. Their semantic similarity may not be high, but their actual fault phenomena are very similar.

[0102] Specifically, this invention can learn a metric function through metric learning, or obtain a metric function through reasonable manual definition.

[0103] The metric function learned by the present invention can include a distance metric algorithm that improves the KNN algorithm, such as the Large Margin Neighbor (LMNN) algorithm.

[0104] Specifically, when obtaining measurement functions through manual definition, fault objects and indicators can be classified and managed hierarchically, using a combination of numbers, letters, and symbols to encode and identify different fault objects and indicators. The classification and grading of fault objects can be constructed according to a CMDB (CMundle Database), as shown in Table 1 below; the classification and grading of indicators can be constructed according to an indicator system (divided into domain and subdomain types), as shown in Table 2 below.

[0105] Table 1 Classification and Grading of Fault Objects

[0106]

[0107]

[0108] Table 2 Classification and Grading Table of Indicators

[0109]

[0110] Specifically, according to the value selection methods in Tables 1 and 2, the present invention can convert abnormal indicators into elements in the corresponding (a, b) format, where 'a' can represent the corresponding value of the faulty object and 'b' can identify the corresponding value of the indicator. For example, when an abnormal indicator is a fault in the service of a physical machine on the platform, the present invention can convert the abnormal indicator into (2.2, 1.1).

[0111] In the case of obtaining the measurement function through manual definition, the present invention can encode each abnormal indicator in the known fault indicator set corresponding to a known fault scenario, so that a single abnormal indicator can be identified by the encoded element, thereby making the known fault indicator set identified by an ordered sequence of encoded elements, such as [(a1,b1),(a2,b2),……(am,bm)].

[0112] Specifically, this invention can train a scene matching model based on an ordered sequence obtained by encoding a known set of fault indicators, until a well-trained scene matching model is obtained. When identifying unknown fault scenarios, each fault indicator in the fault indicator set corresponding to the unknown fault scenario can be encoded and converted to obtain the corresponding ordered sequence [(c1,d1),(c2,d2),……(cn,dn)]. Then, the unknown fault scenario can be identified based on the scene matching model and the ordered sequence.

[0113] The fault handling method proposed in this embodiment can use a scenario matching model to determine a known fault scenario that matches the unknown fault scenario. Then, the loss prevention and control method corresponding to the known fault scenario is used to perform fault recovery processing on the unknown fault scenario, thereby further improving fault handling efficiency and ensuring system operating efficiency.

[0114] based on Figure 1 The method shown in the example proposes a third fault handling method in this embodiment. In this method, based on a first key indicator, it is determined whether a first fault scenario matches at least one known fault scenario in a known fault scenario group, including:

[0115] S301. Determine the third known fault scenario corresponding to the first key indicator; wherein, the third known fault scenario includes at least one predefined scenario indicator;

[0116] The first key indicator can be any one key indicator.

[0117] It should be noted that after determining multiple key indicators, this invention can determine known fault scenarios that match the first fault scenario based on each key indicator. Specifically, in the process of determining known fault scenarios that match the first fault scenario based on the first key indicator, this invention can first determine a third known fault scenario corresponding to the first key indicator, that is, a known fault scenario that includes the first key indicator.

[0118] The scenario indicator can be a known fault indicator from the set of known fault indicators corresponding to the third known fault scenario. It can be understood that the first key indicator can be a scenario indicator included in the third known fault scenario. Optionally, the scenario indicator may also include normal indicators.

[0119] S302. Determine whether there are any scenario indicators other than the first key indicator among the abnormal indicators. If so, proceed to S303; otherwise, prohibit the execution of step S303.

[0120] Specifically, the present invention can determine whether there is a scenario indicator in a third known fault scenario other than the first key indicator among the various abnormal indicators that have appeared in the target system. If so, step S303 can be executed; otherwise, step S303 can be prohibited to avoid unnecessary resource consumption.

[0121] S303. Among the abnormal indicators, determine the occurrence time of each scenario indicator except for the first key indicator;

[0122] Specifically, the present invention can first determine all scenario indicators of the third known fault scenario, excluding the first key indicator, among the various abnormal indicators that have appeared in the target system, and then determine the occurrence time of each scenario indicator.

[0123] S304. Determine the earliest occurrence time from all occurrence times;

[0124] Specifically, this invention can determine the earliest occurrence time from the occurrence times of the known scene indicators.

[0125] S305. The difference between the time of occurrence of the first key indicator and the earliest time of occurrence is determined as the first difference.

[0126] Specifically, the present invention can first determine the occurrence time of the first key indicator, and then subtract the earliest occurrence time from the occurrence time of the first key indicator, and determine the difference obtained by the subtraction as the first difference.

[0127] S306. Based on the first difference, determine whether the first fault scenario matches the third known fault scenario.

[0128] Understandably, if the first difference is not less than zero, it indicates that the scenario indicator of the earliest known third fault scenario in the first fault scenario is not the first critical indicator, and the aforementioned earliest occurrence time is the earliest occurrence time of the scenario indicator of the third known fault scenario in the first fault scenario. Conversely, if the first difference is less than zero, it indicates that the scenario indicator of the earliest known third fault scenario in the first fault scenario is the first critical indicator.

[0129] Specifically, when the present invention determines the scenario indicators of the third known fault scenario starting to appear in the target system, it can start timing to determine whether all scenario indicators of the third known fault scenario appear within a specified time period. If so, it can be determined that the first fault scenario matches the third known fault scenario; otherwise, it can be determined that the first fault scenario does not match the third known fault scenario.

[0130] It should be noted that the present invention can use each key indicator identified from each abnormal indicator as a first key indicator to execute steps S301, S302, S303, S304, S305 and S306 to determine whether there is a matching known fault scenario for the first fault scenario.

[0131] Optionally, step S306 above may include:

[0132] If the first difference is greater than the predefined duration, then it is determined that the first fault scenario does not match the third known fault scenario.

[0133] The predefined duration can be the time window corresponding to the third known fault scenario, i.e. the specified duration mentioned above. The specific duration can be set by technicians according to the actual situation, and this invention does not limit it.

[0134] Specifically, when the first difference is greater than the predefined duration, it can be considered that the target system has not exhibited all scenario indicators of the third known fault scenario within the predefined duration. This indicates that the first fault scenario and the third known fault scenario do not match.

[0135] Optionally, step S306 above may include:

[0136] If the first difference is not less than zero and not greater than the predefined duration, determine whether each abnormal indicator has included all scenario indicators in the third known fault scenario.

[0137] If so, then determine that the first fault scenario matches the third known fault scenario;

[0138] Otherwise, within a predefined duration starting from the earliest occurrence time, check whether all scenario indicators in the third known fault scenario appear in the target system. If so, determine that the first fault scenario matches the third known fault scenario; otherwise, determine that the first fault scenario does not match the third known fault scenario.

[0139] Specifically, when the first difference is not less than zero and not greater than the predefined duration, it can be concluded that the scenario indicator of the earliest known third fault scenario in the first fault scenario is not the first key indicator, and the earliest occurrence time mentioned above is the earliest occurrence time of the scenario indicator of the third known fault scenario in the first fault scenario.

[0140] Optionally, step S306 above may include:

[0141] If the first difference is less than zero, then within a predefined time period starting from the time the first key indicator appears, check whether all scenario indicators in the third known fault scenario appear in the target system. If so, determine that the first fault scenario matches the third known fault scenario; otherwise, determine that the first fault scenario does not match the third known fault scenario.

[0142] Specifically, when the first difference is less than zero, it indicates that the scenario indicator of the earliest known third fault scenario in the first fault scenario is the first key indicator. In this case, if all scenario indicators of the third fault occur in the target system within a predefined timeframe starting from the occurrence of the first key indicator, then the first fault scenario matches the third known fault scenario. If, within a predefined timeframe starting from the occurrence of the first key indicator, not all scenario indicators of the third fault occur in the target system, then the first fault scenario does not match the third known fault scenario.

[0143] Optionally, the third fault handling method described above may further include step S307; wherein:

[0144] S307. If there are no scenario indicators other than the first key indicator among the abnormal indicators, then within a predefined time period starting from the time the first key indicator appears, check whether all scenario indicators in the third known fault scenario appear in the target system. If so, determine that the first fault scenario matches the third known fault scenario; otherwise, determine that the first fault scenario does not match the third known fault scenario.

[0145] It should also be noted that this invention can construct known fault scenarios by using different scenario indicators as root or leaf nodes of a tree according to the tree structure and transmission relationship. The constructed different known fault scenarios can be regarded as a known fault scenario tree. Among them, the loss-stopping processing method can be the conditional triggering action of the leaf node. A path from the root node to the leaf node can correspond to a set of known fault scenarios, that is, a known fault scenario. Furthermore, this invention can merge the component tree and the troubleshooting tree together, and in fact, the troubleshooting tree necessarily contains the known fault tree (except for the conditional triggering action).

[0146] Specifically, this invention presents known fault scenarios in a tree structure, which facilitates data organization, clearly expresses the causal relationship of faults, is easy to interface with troubleshooting trees, and allows the matching of known fault scenarios and the analysis of unknown fault scenarios to be combined.

[0147] To better illustrate the structure of the known fault scenario tree, this invention proposes and combines... Figure 3 Let me introduce it. Figure 3 In this process, the fault indicators of known fault scenarios and unknown fault scenarios are merged. A known fault scenario is only a part of a certain branch and will not cross branches. At this time, all known fault scenario nodes in the branch containing known fault scenario nodes (i.e., root nodes or leaf nodes) and their stop-loss processing methods can form a known fault scenario and its corresponding stop-loss processing method.

[0148] exist Figure 3 In the diagram, a box can identify a fault scenario node, and a circle can identify a loss mitigation plan corresponding to a known fault scenario. Figure 3 Unknown failure scenarios can include "Transaction code - low system success rate", "System launched - within three days", "Exchange in AP-MQ - queue handle not released", "Exchange in AP-WebLogic - connection pool full", "Exchange using Oracle - database table - many row locks", and "Exchange using Oracle - all indexes are invalid". Figure 3The known fault scenarios can include "Oracle used by the exchange - all indexes are invalid", "Oracle used by the exchange - high AAS", "High CPU wait IO in the exchange's AP", "Oracle machine used by the exchange - high CPU utilization" and "SAN storage used by Oracle - IOPS exceeded". Figure 3 Mitigation measures for known faulty nodes can include "version rollback", "service restart", "killing sessions that generate row locks" and "index rebuilding".

[0149] When an alarm or anomaly triggers troubleshooting (not necessarily the root node), this invention can prioritize searching among known fault scenario nodes in the tree, and simultaneously begin updating unknown fault scenario nodes without affecting efficiency. Following... Figure 1 If the method shown can match a known fault scenario, it will directly output the corresponding loss-stopping method and continue the root cause analysis of the troubleshooting tree; if it cannot match, it will continue the root cause analysis of the troubleshooting tree and use the combination of abnormal nodes (i.e., abnormal indicators) in the troubleshooting tree to match unknown fault scenarios and recommend loss-stopping methods.

[0150] It should also be noted that, based on Figure 3 As shown, this invention can use Spearman Rank Correlation (SRC) to determine the similarity between (a1,b1) and (c1,d1). Since the fault sequence values ​​do not consider the original numerical values ​​of the data, but rather rank the data according to a certain method, SRC is very suitable. Furthermore, since the values ​​of m and n are likely to be different, and considering that objects and indicators not appearing in the fault scenario records are generally irrelevant to loss mitigation, after ranking the anomalies provided by the fault scenario records and the troubleshooting tree, they are aligned according to object domains and indicator domains respectively. SRC is only calculated for these aligned sequences, and a weighted average is taken between the two results.

[0151] It should also be noted that, based on Figure 3 As shown, when determining the similarity between (a1,b1) and (c1,d1), this invention can also use the shorter sequence as a sliding window, sliding along the longer sequence by one unit length. Each sliding operation yields a sliding similarity, and finally, the weighted average of all sliding similarities is calculated to obtain the final similarity. Since the fault sequences are ordered, i.e., gradually approaching the root cause of the fault, and the similarity of the root cause contributes significantly to the stop-loss recommendation, the sliding should start from the end after alignment, and the weights should be set to decrease sequentially (the rule is to decrease by a factor of 2 each time, with a cumulative sum of 100%).

[0152] It should also be noted that during the matching process of unknown fault scenarios, the present invention can accumulate a certain amount of instantiated matching data. At this time, the matching data can be further used to train the above-mentioned scenario matching model, thereby improving the utilization rate of the data.

[0153] The fault handling method provided in this embodiment can realize the matching process between the first fault scenario and the known fault scenario, thereby further improving the fault handling efficiency and ensuring the system operating efficiency.

[0154] and Figure 1 The steps shown correspond to those in the example below, such as... Figure 4 As shown, this embodiment proposes a first fault handling device. This device may include: a first obtaining unit 101, a first identifying unit 102, a first determining unit 103, a second determining unit 104, and a first processing unit 105; wherein:

[0155] The first obtaining unit 101 is used to obtain at least one abnormal indicator that has appeared in the target system in response to a fault diagnosis command.

[0156] The first identification unit 102 is used to identify at least one key indicator from each abnormal indicator;

[0157] The first determining unit 103 is used to determine whether the current first fault scenario of the target system matches at least one known fault scenario in the known fault scenario group based on the identified key indicators respectively. If so, the second determining unit 104 is triggered.

[0158] The second determining unit 104 is used to determine a first known fault scenario that matches the first fault scenario from the known fault scenario group;

[0159] The first processing unit 105 is used to perform fault handling using a loss mitigation method corresponding to the first known fault scenario.

[0160] It should be noted that the first obtaining unit 101, the first identifying unit 102, the first determining unit 103, the second determining unit 104, and the first processing unit 105 can be referred to respectively. Figure 1 The relevant explanations of steps S101, S102, S103, S104 and S105 are not repeated here.

[0161] Optionally, the device may further include: a third determining unit, a fourth determining unit, and a second processing unit;

[0162] The third determining unit is used to determine the first fault scenario as an unknown fault scenario if the first fault scenario does not match any known fault scenario in the known fault scenario group.

[0163] The fourth determining unit is used to use the trained scene matching model to determine the second known fault scene that matches the first fault scene from the known fault scene group;

[0164] The second processing unit is used to handle faults using a loss mitigation method corresponding to the second known fault scenario.

[0165] Optionally, based on the first key indicator, determine whether the first fault scenario matches at least one known fault scenario in the known fault scenario group, set as follows:

[0166] Identify the third known fault scenario corresponding to the first key indicator; wherein the third known fault scenario includes at least one predefined scenario indicator;

[0167] Determine whether there are any scenario-related indicators other than the primary key indicator among the abnormal indicators;

[0168] If so, then among the abnormal indicators, determine the occurrence time of each scenario indicator other than the first key indicator, and determine the earliest occurrence time from among the occurrence times.

[0169] The first difference is determined by subtracting the earliest occurrence time from the occurrence time of the first key indicator.

[0170] Based on the first difference, determine whether the first fault scenario matches the third known fault scenario.

[0171] Optionally, based on the first difference, it is determined whether the first fault scenario matches the third known fault scenario, set as follows:

[0172] If the first difference is greater than the predefined duration, then it is determined that the first fault scenario does not match the third known fault scenario.

[0173] Optionally, based on the first difference, it is determined whether the first fault scenario matches the third known fault scenario, set as follows:

[0174] If the first difference is not less than zero and not greater than the predefined duration, determine whether each abnormal indicator has included all scenario indicators in the third known fault scenario.

[0175] If so, then determine that the first fault scenario matches the third known fault scenario;

[0176] Otherwise, within a predefined duration starting from the earliest occurrence time, check whether all scenario indicators in the third known fault scenario appear in the target system. If so, determine that the first fault scenario matches the third known fault scenario; otherwise, determine that the first fault scenario does not match the third known fault scenario.

[0177] Optionally, based on the first difference, it is determined whether the first fault scenario matches the third known fault scenario, set as follows:

[0178] If the first difference is less than zero, then within a predefined time period starting from the time the first key indicator appears, check whether all scenario indicators in the third known fault scenario appear in the target system. If so, determine that the first fault scenario matches the third known fault scenario; otherwise, determine that the first fault scenario does not match the third known fault scenario.

[0179] Optionally, the device further includes: a detection unit, a fifth determining unit, and a sixth determining unit; wherein:

[0180] The detection unit is used to detect whether all scenario indicators in the third known fault scenario appear in the target system within a predefined time period starting from the time the first key indicator appears if there are no scenario indicators other than the first key indicator among the abnormal indicators. If so, the fifth determination unit is triggered; otherwise, the sixth determination unit is triggered.

[0181] The fifth determining unit is used to determine whether the first fault scenario matches the third known fault scenario;

[0182] The sixth determining unit is used to determine that the first fault scenario does not match the third known fault scenario.

[0183] The fault handling device proposed in this embodiment can effectively improve fault handling efficiency and ensure system operating efficiency.

[0184] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0185] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A fault handling method, characterized in that, include: In response to a fault diagnosis command, obtain at least one abnormal indicator that has appeared in the target system. Identify at least one key indicator from each of the aforementioned abnormal indicators; Based on each of the identified key indicators, determine whether the current first fault scenario of the target system matches at least one known fault scenario in the known fault scenario group; If so, a first known fault scenario matching the first fault scenario is determined from the known fault scenario group, and the fault is handled using the loss-stopping handling method corresponding to the first known fault scenario. Specifically, determining whether the first fault scenario matches at least one known fault scenario in the known fault scenario group based on the first key indicator includes: Determine the third known fault scenario corresponding to the first key indicator; wherein, the third known fault scenario includes at least one predefined scenario indicator; Determine whether any of the aforementioned abnormal indicators are other than the first key indicator; If so, then among the abnormal indicators, determine the occurrence time of each of the scenario indicators other than the first key indicator, and determine the earliest occurrence time from among the occurrence times; The difference obtained by subtracting the earliest occurrence time from the occurrence time of the first key indicator is determined as the first difference; Based on the first difference, it is determined whether the first fault scenario matches the third known fault scenario.

2. The fault handling method according to claim 1, characterized in that, The method further includes: If the first fault scenario does not match any known fault scenario in the known fault scenario group, then the first fault scenario is determined to be an unknown fault scenario. Using a trained scene matching model, a second known fault scene that matches the first fault scene is determined from the known fault scene group, and the fault is handled using the loss prevention handling method corresponding to the second known fault scene.

3. The fault handling method according to claim 1, characterized in that, The step of determining whether the first fault scenario matches the third known fault scenario based on the first difference includes: If the first difference is greater than a predefined duration, then it is determined that the first fault scenario does not match the third known fault scenario.

4. The fault handling method according to claim 1, characterized in that, The step of determining whether the first fault scenario matches the third known fault scenario based on the first difference includes: If the first difference is not less than zero and not greater than a predefined duration, determine whether each of the abnormal indicators has included all the scenario indicators in the third known fault scenario; If so, then the first fault scenario is determined to match the third known fault scenario; Otherwise, within a predefined duration starting from the earliest occurrence time, it is detected in the target system whether all the scenario indicators in the third known fault scenario appear. If so, it is determined that the first fault scenario matches the third known fault scenario; otherwise, it is determined that the first fault scenario does not match the third known fault scenario.

5. The fault handling method according to claim 1, characterized in that, The step of determining whether the first fault scenario matches the third known fault scenario based on the first difference includes: If the first difference is less than zero, then within a predefined time period starting from the time the first key indicator appears, it is detected in the target system whether all the scenario indicators in the third known fault scenario appear. If so, it is determined that the first fault scenario matches the third known fault scenario; otherwise, it is determined that the first fault scenario does not match the third known fault scenario.

6. The fault handling method according to claim 1, characterized in that, The method further includes: If none of the scenario indicators other than the first key indicator exist among the abnormal indicators, then within a predefined time period starting from the occurrence time of the first key indicator, it is detected in the target system whether all the scenario indicators in the third known fault scenario appear. If so, it is determined that the first fault scenario matches the third known fault scenario; otherwise, it is determined that the first fault scenario does not match the third known fault scenario.

7. A fault handling device, characterized in that, include: The system comprises a first obtaining unit, a first identifying unit, a first determining unit, a second determining unit, and a first processing unit; wherein: The first obtaining unit is configured to obtain at least one abnormal indicator that has appeared in the target system in response to a fault diagnosis command; The first identification unit is used to identify at least one key indicator from each of the abnormal indicators; The first determining unit is used to determine, based on each of the identified key indicators, whether the current first fault scenario of the target system matches at least one known fault scenario in the known fault scenario group; if so, the second determining unit is triggered. The second determining unit is used to determine a first known fault scenario that matches the first fault scenario from the known fault scenario group; The first processing unit is used to perform fault handling using a loss mitigation method corresponding to the first known fault scenario; Specifically, determining whether the first fault scenario matches at least one known fault scenario in the known fault scenario group based on the first key indicator is set as follows: Determine the third known fault scenario corresponding to the first key indicator; wherein, the third known fault scenario includes at least one predefined scenario indicator; Determine whether any of the aforementioned abnormal indicators are other than the first key indicator; If so, then among the abnormal indicators, determine the occurrence time of each of the scenario indicators other than the first key indicator, and determine the earliest occurrence time from among the occurrence times; The difference obtained by subtracting the earliest occurrence time from the occurrence time of the first key indicator is determined as the first difference; Based on the first difference, it is determined whether the first fault scenario matches the third known fault scenario.

8. The fault handling device according to claim 7, characterized in that, The device further includes: a third determining unit, a fourth determining unit, and a second processing unit; The third determining unit is configured to determine the first fault scenario as an unknown fault scenario if the first fault scenario does not match any known fault scenario in the known fault scenario group. The fourth determining unit is used to use a trained scene matching model to determine a second known fault scene that matches the first fault scene from the known fault scene group. The second processing unit is used to perform fault handling using a loss mitigation method corresponding to the second known fault scenario.

Citation Information

Patent Citations

  • Fault solution method and device

    CN111859047A

  • Fault handling method and device, medium and equipment

    CN112446511A