A fault processing method, device, system and computer readable storage medium

By automating the adjustment of fault handling priorities and methods in the computer system, the problems of poor timeliness and accuracy of manual inspections have been solved, achieving efficient and reliable fault handling and improving system reliability and stability.

CN122309202APending Publication Date: 2026-06-30CONTEMPORARY AMPEREX FUTURE ENERGY RES INST (SHANGHAI) LTD +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CONTEMPORARY AMPEREX FUTURE ENERGY RES INST (SHANGHAI) LTD
Filing Date
2024-12-31
Publication Date
2026-06-30

AI Technical Summary

Technical Problem

In existing technologies, computer system fault handling relies on manual inspection, which is time-consuming and inaccurate, affecting production efficiency and safe operation.

Method used

By acquiring relevant data and adjusting processing priorities when a target fault is detected, computer system faults can be handled automatically. The mapping information and fault association information in the configuration information can be used to dynamically adjust the processing priority and method of the fault.

Benefits of technology

It improves the efficiency and reliability of fault handling, reduces the workload of maintenance personnel, ensures that important or urgent faults are handled in a timely manner, and enhances the reliability and operational stability of computer systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122309202A_ABST
    Figure CN122309202A_ABST
Patent Text Reader

Abstract

This application discloses a fault handling method, apparatus, system, and computer-readable storage medium. The method, applied to a computer system, includes: upon detecting a target fault, acquiring target data related to the target fault and obtaining an initial processing priority for the target fault from configuration information; adjusting the initial processing priority of the target fault based on the target data; and processing the target fault according to the adjusted processing priority. Through these methods, this application ensures that relatively important or urgent target faults are handled preferentially and promptly, improving the efficiency and reliability of fault handling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of fault handling technology, and in particular to a fault handling method, apparatus, system and computer-readable storage medium. Background Technology

[0002] With the rapid rise of concepts such as Industry 4.0 and smart manufacturing, the reliability and maintainability of computer systems have become critical. Complex computer systems (such as embedded systems) are widely used in various fields such as industrial automation, transportation, and healthcare, and computer system failures have a significant impact on production efficiency and safe operation.

[0003] Currently, most computer system fault handling relies on manual inspections and periodic maintenance; however, manual fault handling is inefficient and inaccurate. Summary of the Invention

[0004] This application provides at least one fault handling method, apparatus, system, and computer-readable storage medium.

[0005] The first aspect of this application provides a fault handling method applied to a computer system, comprising: upon detecting a target fault, acquiring target data related to the target fault, and acquiring an initial processing priority of the target fault from configuration information; adjusting the initial processing priority of the target fault based on the target data; and processing the target fault according to the adjusted processing priority of the target fault.

[0006] Therefore, by adjusting the processing priority of the target fault, it can be determined whether the target fault is urgent or important. Thus, processing the target fault according to the adjusted processing priority can ensure that relatively important or relatively urgent target faults are processed in a timely manner, improving the efficiency and reliability of fault handling, thereby improving the reliability and operational stability of the computer system.

[0007] It automatically handles target faults upon detection, enabling rapid response and processing, making fault handling more efficient and reliable, and reducing the workload of maintenance personnel.

[0008] The target data includes at least one of the following: system parameters of the computer system when or after the target fault occurs, fault information affecting the target fault, and fault association information. The fault association information indicates whether there is a pre-defined correlation between the target fault and other faults, and the fault information affecting the target fault can characterize the impact on the performance of the computer system.

[0009] Therefore, the target data can be set flexibly.

[0010] The system parameters include at least one of system status parameters and system performance parameters. The system status parameters include at least one of the following: CPU utilization, memory utilization, disk read / write speed, and data transmission speed. The system performance parameters include at least one of the following: system load, response time, and fault frequency. The fault-affecting information includes at least one of the following: historical fault handling results of the target fault and fault occurrence frequency. The target fault has a pre-defined correlation with other faults: the occurrence of the target fault will lead to the occurrence of other faults.

[0011] Therefore, system status parameters, system performance parameters, and information affecting faults can be flexibly set.

[0012] The target data includes system parameters, which include at least two system state parameters. Based on the target data, the initial processing priority of the target fault is adjusted, including: obtaining the weights corresponding to each system state parameter; weighting each system state parameter based on its weight to obtain a fault score for the target fault; and using the fault score to adjust the initial processing priority of the target fault.

[0013] Therefore, the fault score of the target fault is obtained by weighting each system state parameter according to its corresponding weight, that is, it fully considers the influence of different system state parameters; therefore, the processing priority of the target fault determined based on the fault score is accurate and reasonable.

[0014] Before obtaining the fault score of the target fault by weighting each system state parameter based on the weights corresponding to each system state parameter, the fault handling method further includes: normalizing each system state parameter; and / or adjusting the initial handling priority of the target fault using the fault score, including: obtaining a reference handling priority corresponding to the fault score; adjusting the handling priority of the target fault to the reference handling priority in response to the reference handling priority being higher than the current handling priority of the target fault, or increasing the handling priority of the target fault by a preset first adjustment step size; and / or decreasing the handling priority of the target fault by a preset second adjustment step size in response to the reference handling priority being lower than or equal to the current handling priority of the target fault.

[0015] Therefore, by normalizing the system state parameters, it is ensured that system state parameters with different value ranges can be compared and calculated on the same scale. When the reference processing priority corresponding to the fault score of the target fault is higher than the current processing priority of the target fault, there are two adjustment methods: one is to adjust the processing priority of the target fault directly to the reference processing priority in one step, and the other is to gradually increase the processing priority of the target fault; when the reference processing priority corresponding to the fault score of the target fault is lower than the current processing priority of the target fault, the processing priority of the target fault is gradually decreased.

[0016] The target data includes system parameters, which include system performance parameters. Based on the target data, the initial processing priority of the target fault is adjusted, including: in response to the system performance parameters exceeding the parameter range and the target fault being a preset associated fault of the system performance parameters, the initial processing priority of the target fault is increased.

[0017] Therefore, it is possible to prioritize and promptly address target faults related to abnormal system parameters.

[0018] The target data includes information affecting the fault; based on the target data, the initial processing priority of the target fault is adjusted, including: in response to the information affecting the fault meeting preset conditions, the initial processing priority of the target fault is increased, wherein, when the information affecting the fault includes historical fault processing results, the preset conditions include the number of historical fault processing results that are failures reaching a preset number threshold within the first historical time period, and when the information affecting the fault includes the fault occurrence frequency, the preset conditions include the fault occurrence frequency being higher than a preset frequency threshold.

[0019] Therefore, for target faults that have been pending for a long time without resolution, their initial processing priority will be increased to ensure timely handling and resolution. Furthermore, even if automated processing fails again, the increased initial processing priority ensures that it is promptly reported to maintenance personnel for further processing, guaranteeing timely resolution. Similarly, for target faults that occur frequently within a short period, their initial processing priority will also be increased to ensure timely handling and resolution.

[0020] The target data includes fault association information; based on the target data, the initial processing priority of the target fault is adjusted, including: in response to the fault association information indicating that the target fault has a preset correlation with other faults, the initial processing priority of the target fault is increased.

[0021] Therefore, by increasing the initial processing priority of the target fault, it is possible to ensure that the target fault, as the root cause, is processed in a timely and prioritized manner. This avoids the occurrence of other faults due to the failure to process the target fault as the root cause in a timely and prioritized manner, thereby avoiding the occurrence of a chain reaction, reducing the impact on the computer system, and ensuring the reliability and operational stability of the computer system.

[0022] The target data includes fault association information, whereby the target fault exhibits a pre-defined correlation with other faults: the occurrence of the target fault leads to the occurrence of other faults. The steps for determining the fault association information of the target fault include: generating several first transactions using historical fault data of the computer system; wherein the first transactions include several historical faults occurring within a second historical time period, and the target fault exists within the several first transactions; performing at least one statistical analysis on the several first transactions to obtain at least one correlation degree characterization value between the target fault and other faults, each correlation degree characterization value characterizing whether the occurrence of the target fault leads to the occurrence of other faults; and determining whether there is a pre-defined correlation between the target fault and other faults based on at least one correlation degree characterization value.

[0023] Therefore, since the correlation degree characterization value represents whether the occurrence of a target fault will lead to the occurrence of other faults, when the correlation degree characterization value represents that the occurrence of a target fault will lead to the occurrence of other faults, it can be determined that there is a pre-defined correlation between the target fault and other faults.

[0024] The method involves performing at least one statistical analysis on several first transactions to obtain at least one correlation characterization value between the target fault and other faults, including: selecting a first transaction containing the target fault from the several first transactions as a second transaction; calculating the proportion of other faults occurring in the second transaction as the confidence level between the target fault and other faults; using the confidence level as a correlation characterization value; and / or calculating the proportion of combinations of the target fault and other faults occurring in the several first transactions as the support level between the target fault and other faults, and using the ratio between the confidence level and the support level as a correlation characterization value.

[0025] Therefore, the correlation characterization value between the target fault and other faults can be flexibly set.

[0026] Specifically, determining whether there is a pre-defined correlation between the target fault and other faults based on at least one correlation degree characterization value includes: in response to each correlation degree characterization value indicating that the occurrence of the target fault will lead to the occurrence of other faults, determining that there is a pre-defined correlation between the target fault and other faults.

[0027] Therefore, when each correlation value indicates that the occurrence of the target fault will lead to the occurrence of other faults, it can be determined that there is a pre-defined correlation between the target fault and other faults.

[0028] The process of obtaining the initial processing priority of the target fault from the configuration information includes: obtaining the first fault identifier of the target fault; querying the processing priority corresponding to the first fault identifier from the configuration information as the initial processing priority of the target fault. The configuration information contains mapping information for different faults of the computer system, and the mapping information includes the first fault identifier of the fault and the corresponding processing priority.

[0029] Therefore, the configuration information is similar to a comprehensive database of computer systems, containing statistical mapping information of different faults in the computer system.

[0030] The fault handling method further includes, before handling the target fault according to the adjusted processing priority, retrieving the corresponding handling method from the configuration information; and handling the target fault according to the adjusted processing priority, including handling the target fault according to the adjusted processing priority.

[0031] Therefore, the configuration information includes preset handling methods for each fault. By utilizing these preset handling methods for the target fault, the target fault can be processed automatically without the need for analysis and determination of the handling method. This enables rapid response and handling of target faults, making fault handling more efficient and reliable, and reducing the workload of maintenance personnel.

[0032] The processing of the target fault includes: in response to the processing method being manual processing, feeding back the first fault-related information of the target fault to the host computer so that the host computer can display the first fault-related information through the first fault pop-up window; and in response to the processing method being automatic processing, processing the target fault according to the automatic processing method.

[0033] Therefore, the system will report the first fault-related information of the target fault to the host computer. The host computer will automatically pop up a fault pop-up window, displaying the first fault-related information of the target fault, to promptly notify maintenance personnel. Because the fault pop-up window displays the first fault-related information, it facilitates the resolution of the target fault by maintenance personnel. If the handling method for the target fault, as determined by the configuration information, is automatic, the target fault will be handled automatically, achieving rapid response and handling of the target fault. This makes the handling of the target fault more efficient and reliable, reducing the workload of maintenance personnel.

[0034] The fault handling method also includes: in response to the failure to handle the target fault according to the automatic handling method, feeding back the second fault-related information of the target fault to the host computer.

[0035] The configuration information includes mapping information for different computer system faults. The mapping information includes the first fault identifier of the fault and the corresponding handling method. The handling method corresponding to the target fault is obtained by querying the configuration information, including: obtaining the first fault identifier of the target fault; and querying the handling method corresponding to the first fault identifier from the configuration information as the handling method corresponding to the target fault.

[0036] Therefore, after the automatic handling of the target fault fails, the system will report the second fault-related information of the target fault to the host computer. The host computer will automatically pop up a fault pop-up window and display the second fault-related information of the target fault in the fault pop-up window, so as to notify the operation and maintenance personnel in a timely manner. Since the fault pop-up window can display the second fault-related information of the target fault, it is convenient for the operation and maintenance personnel to resolve the target fault.

[0037] The first fault identifier consists of the device type identifier, board type identifier, module type identifier and second fault identifier of the target fault, and the second fault identifier is the identifier corresponding to the type of the target fault.

[0038] Therefore, by establishing a structured first fault identification system, the fault description of the target fault can be more detailed, enabling maintenance personnel to quickly and accurately locate the target fault.

[0039] The fault handling method further includes at least one of the following steps: updating the fault status of the target fault based on the handling result of the target fault; recording the handling-related information of the target fault; wherein the handling-related information includes at least one of the following: handling time, handling method, and handling result; in response to detecting the target fault, generating a fault event of the target fault and feeding the fault event back to the host computer so that the host computer displays the fault event of the target fault through a second fault pop-up window; in response to a query request for the target fault issued by the host computer, feeding back the fault-related information of the target fault to the host computer; wherein the fault event includes at least one of the first fault identifier and mapping information of the target fault, and the mapping information includes at least one of the following: the first fault identifier of the target fault, the fault type of the target fault, the fault name of the target fault, the handling priority corresponding to the target fault, and the handling method corresponding to the target fault.

[0040] A second aspect of this application provides a fault handling apparatus, which includes an acquisition module, a determination module, and a processing module. The acquisition module is used to acquire target data related to the target fault when a target fault is detected, and to acquire the initial processing priority of the target fault from configuration information. The adjustment module is used to adjust the initial processing priority of the target fault based on the target data. The processing module is used to process the target fault according to the adjusted processing priority of the target fault.

[0041] A third aspect of this application provides an electronic device including a processor and a memory, the memory storing program instructions, and the processor executing the program instructions to implement the above-described fault handling method.

[0042] The fourth aspect of this application provides a fault handling system, which includes a host computer and a slave computer, wherein the slave computer is the aforementioned electronic device.

[0043] The fifth aspect of this application provides a computer-readable storage medium for storing program instructions that can be executed to implement the above-described fault handling method.

[0044] The above technical solution can determine whether a target fault is urgent or important by adjusting the processing priority of the target fault. Therefore, processing the target fault according to the adjusted processing priority can ensure that relatively important or relatively urgent target faults are processed in a timely manner, thereby improving the efficiency and reliability of fault handling and thus improving the reliability and operational stability of the computer system.

[0045] Upon detecting a target fault, the system automatically handles the fault, enabling rapid response and resolution, making fault handling more efficient and reliable, and reducing the workload of maintenance personnel. Attached Figure Description

[0046] Figure 1 This is a flowchart illustrating an embodiment of the fault handling method provided in this application;

[0047] Figure 2 This is a schematic diagram of the framework structure of an embodiment of the fault handling system provided in this application;

[0048] Figure 3 This is a structural schematic diagram of an embodiment of the first fault identifier provided in this application;

[0049] Figure 4 This is a flowchart illustrating an embodiment of the fault handling process provided in this application;

[0050] Figure 5 yes Figure 1The flowchart of step S13 shown is a schematic diagram of one embodiment.

[0051] Figure 6 This is a flowchart illustrating an embodiment of the method for determining fault association information of a target fault provided in this application;

[0052] Figure 7 yes Figure 6 The flowchart of step S62 shown is a schematic diagram of one embodiment.

[0053] Figure 8 This is a schematic diagram of the structure of an embodiment of the fault handling device provided in this application;

[0054] Figure 9 This is a schematic diagram of the structure of an embodiment of the electronic device provided in this application;

[0055] Figure 10 This is a schematic diagram of an embodiment of the computer-readable storage medium provided in this application. Detailed Implementation

[0056] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0057] In the following description, specific details such as particular system architectures, interfaces, and technologies are presented for illustrative purposes rather than for limiting purposes, in order to provide a thorough understanding of this application.

[0058] In this document, the term "and / or" is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this document means two or more. Moreover, the term "at least one" in this document means any combination of at least two of any one or more of a plurality of objects. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0059] Please see Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the fault handling method provided in this application. It should be noted that if substantially the same result is achieved, this embodiment does not necessarily replace it with a similar method. Figure 1 The illustrated process sequence is limited. For example... Figure 1 As shown, the fault handling method is applied to a computer system. This embodiment includes:

[0060] Step S11: If a target fault is detected, obtain target data related to the target fault, and obtain the initial processing priority of the target fault from the configuration information.

[0061] In this embodiment, a target fault in the computer system is detected. That is, real-time fault monitoring of the computer system is possible, allowing for timely intervention when a fault is detected, thus improving fault handling efficiency and reducing the impact and losses caused by the fault.

[0062] In one embodiment, the computer system may be an embedded system. An embedded system is a dedicated computer system designed to perform specific functions or tasks, typically embedded as a component of other devices or systems. These systems are widely used in everyday life, such as smartphones, home appliances, medical devices, and industrial applications.

[0063] Of course, in other implementations, the computer system may also be a distributed system, a real-time system, a network system, etc., and this is not limited here.

[0064] In one implementation, such as Figure 2 As shown, Figure 2 This is a schematic diagram of the framework structure of an embodiment of the fault handling system provided in this application. The fault handling system includes a host computer and a slave computer, which communicate via a TCP connection. The slave computer includes a fault triggering module. Figure 2 The "fault triggering" module detects whether a target fault has occurred in the computer system.

[0065] In one implementation, in response to the detection of a target fault, a fault event for the target fault is generated and fed back to the host computer, so that the host computer displays the fault event for the target fault through a second fault pop-up window. That is, upon detection of a target fault, the generated fault event for the target fault is fed back to the host computer, which automatically pops up a fault pop-up window and displays the fault event for the target fault. This provides a clear view of the target fault occurring in the computer system to maintenance personnel, enabling them to promptly become aware of the target fault occurring in the computer system.

[0066] The fault event includes at least one of the first fault identifier of the target fault and mapping information. The mapping information includes at least one of the following: the first fault identifier of the target fault, the fault type of the target fault, the fault name of the target fault, the processing priority corresponding to the target fault, and the processing method corresponding to the target fault.

[0067] In one specific implementation, the fault event includes a first fault identifier and mapping information of the target fault. Therefore, displaying the fault event of the target fault is also displaying the first fault identifier and mapping information of the target fault. Thus, displaying the fault event that includes the first fault identifier and mapping information of the target fault, and centrally displaying the first fault identifier and mapping information of the target fault, improves the readability of the fault event of the target fault, making it easier for maintenance personnel to quickly and accurately identify and understand the fault.

[0068] In one specific implementation, the fault event includes mapping information, which includes a first fault identifier of the target fault, the fault type of the target fault, the fault name of the target fault, the processing priority of the target fault, and the processing method of the target fault. By displaying the first fault identifier, the fault type, the fault name, the processing priority, and the processing method of the target fault, detailed information about the target fault is presented to maintenance personnel, facilitating quick and accurate identification and understanding of the fault.

[0069] In one specific implementation, such as Figure 3 As shown, Figure 3 This is a schematic diagram of a first fault identifier embodiment provided in this application. The first fault identifier is identified by the device type where the target fault occurs. Figure 3 "Device type" and board type identifier (in the text) Figure 3 "Board type" and module type identifier (in the text) Figure 3 "Module type" and second fault identifier ( Figure 3 The first fault identifier consists of the "fault code" and the second fault identifier is the identifier corresponding to the type of the target fault. By establishing a structured first fault identifier system, the fault description of the target fault becomes more detailed, enabling maintenance personnel to quickly and accurately locate the target fault.

[0070] In one specific implementation, such as Figure 2 As shown, the lower-level computer includes a fault triggering module. In response to detecting a target fault in the computer system, the fault triggering module generates a fault event of the target fault and feeds the fault event back to the upper-level computer so that the upper-level computer can display the fault event of the target fault through a second fault pop-up window.

[0071] In addition, in this embodiment, target data related to the target fault will be acquired.

[0072] In one embodiment, the target data includes at least one of the following: system parameters of the computer system when or after the target fault occurs, impact fault information of the target fault, and fault association information. The fault association information indicates whether there is a predetermined correlation between the target fault and other faults, and the impact fault information can characterize the impact on the performance of the computer system. The target data includes multiple types; therefore, by acquiring different types of target data, the initial processing priority of the target fault can be dynamically adjusted subsequently using these different types of target data. Furthermore, the target data used for dynamically adjusting the initial processing priority of the target fault can be flexibly selected.

[0073] In one specific implementation, the system parameters include at least one of system status parameters and system performance parameters. The system status parameters include at least one of the following: CPU utilization, memory utilization, disk read / write speed, and data transmission speed. The system performance parameters include at least one of the following: system load, response time, and fault frequency. The fault-affecting information includes at least one of the historical fault handling results and fault occurrence frequency of the target fault. The target fault exhibits a preset correlation with other faults: the occurrence of the target fault leads to the occurrence of other faults. Since the system status parameters, system performance parameters, and fault-affecting information are of various types, by acquiring different types of system status parameters, system performance parameters, and fault-affecting information, the initial processing priority of the target fault can be dynamically adjusted subsequently. Furthermore, the system status parameters, system performance parameters, and fault-affecting information used for dynamically adjusting the initial processing priority of the target fault can be flexibly selected.

[0074] In one implementation, obtaining the initial processing priority of the target fault from the configuration information specifically involves: obtaining the first fault identifier of the target fault; querying the configuration information to find the processing priority corresponding to the first fault identifier, which is then used as the initial processing priority of the target fault. In other words, the configuration information is similar to a comprehensive database in a computer system, and the first fault identifier is the fault identifier corresponding to the target fault; therefore, obtaining the processing priority corresponding to the first fault identifier from the configuration information is the initial processing priority of the target fault corresponding to the first fault identifier.

[0075] Step S12: Based on the target data, adjust the initial processing priority of the target fault.

[0076] In this embodiment, the initial processing priority of target faults is adjusted based on target data. In other words, the initial processing priority of target faults occurring in the computer system can be dynamically adjusted based on target data related to those faults, ensuring that relatively important or urgent target faults are handled preferentially and promptly. This improves the efficiency and reliability of fault handling, thereby enhancing the reliability and operational stability of the computer system.

[0077] It should be noted that adjusting the initial processing priority of the target fault can mean either increasing or decreasing the initial processing priority of the target fault. Alternatively, adjusting the initial processing priority of the target fault can mean increasing or decreasing the initial processing priority of the target fault by a certain step size, etc.

[0078] In one implementation, such as Figure 2 As shown, the lower-level machine also includes a dynamic priority adjustment module. Figure 2 The "dynamic priority adjustment" module adjusts the initial processing priority of the target fault based on the target data.

[0079] Step S13: Process the target fault according to the adjusted processing priority of the target fault.

[0080] In this implementation, the target fault is processed according to its adjusted processing priority. By determining the adjusted processing priority, the urgency or importance of the target fault can be identified. Therefore, processing the target fault according to its adjusted priority ensures that relatively important or urgent faults are handled preferentially and promptly, improving the efficiency and reliability of fault handling, thereby enhancing the reliability and operational stability of the computer system. Furthermore, upon detection of a target fault, the processing is automated, enabling rapid response and handling, making fault handling more efficient and reliable, and reducing the workload of maintenance personnel.

[0081] In one implementation, such as Figure 2 As shown, the lower-level machine also includes a fault handling module. Figure 2 The "fault handling" module processes the target fault according to the adjusted processing priority of the target fault.

[0082] In one embodiment, before processing the target fault according to the adjusted processing priority of the target fault, the processing method corresponding to the target fault is obtained from the configuration information; at this time, the target fault is processed according to the adjusted processing priority of the target fault, specifically: the target fault is processed according to the adjusted processing priority of the target fault using the processing method.

[0083] In other words, the configuration information contains preset handling methods for each fault. By utilizing these preset handling methods for the target fault, the target fault can be processed automatically without the need for analysis and determination of the handling method. This enables rapid response and handling of target faults, making fault handling more efficient and reliable, and reducing the workload of maintenance personnel.

[0084] In one specific implementation, the target fault is processed using a processing method, specifically: in response to the processing method being manual processing, first fault-related information of the target fault is fed back to the host computer, so that the host computer displays the first fault-related information through a first fault pop-up window; in response to the processing method being automatic processing, the target fault is processed according to the automatic processing method. That is to say, as... Figure 4 As shown, Figure 4 This is a flowchart illustrating an embodiment of the fault handling process provided in this application. When the handling method corresponding to the target fault, as determined by the configuration information, is manual, the system will send the first fault-related information of the target fault to the host computer. The host computer will automatically pop up a fault pop-up window, displaying the first fault-related information of the target fault, to promptly notify maintenance personnel. Since the fault pop-up window can display the first fault-related information of the target fault, it facilitates the maintenance personnel in resolving the target fault. When the handling method corresponding to the target fault, as determined by the configuration information, is automatic, the target fault will be handled according to the automatic handling method, achieving rapid response and handling of the target fault. This makes the handling of the target fault more efficient and reliable, reducing the workload of maintenance personnel.

[0085] The automatic handling methods can include stopping the module to which the target fault belongs, restarting the module to which the target fault belongs, restarting the computer system, stopping the computer system, ignoring, etc.

[0086] It should be noted that when the handling method for the target fault obtained from the configuration information is manual handling, the information should be promptly fed back to the operation and maintenance personnel through the host computer so that the operation and maintenance personnel can handle the target fault in a timely manner. This can also be understood as a rapid response and handling of the target fault.

[0087] In one specific implementation, in response to the failure of automatic processing of the target fault, a second fault-related information of the target fault is fed back to the host computer. That is, after the automatic processing of the target fault fails, the second fault-related information of the target fault is fed back to the host computer, which will automatically pop up a fault pop-up window and display the second fault-related information of the target fault in the fault pop-up window, so as to notify the operation and maintenance personnel in a timely manner. Since the fault pop-up window can display the second fault-related information of the target fault, it is convenient for the operation and maintenance personnel to resolve the target fault.

[0088] In one specific implementation, the configuration information includes mapping information for different computer system faults. The mapping information includes a first fault identifier and a corresponding handling method. In this case, retrieving the handling method corresponding to the target fault from the configuration information includes: obtaining the first fault identifier of the target fault; and retrieving the handling method corresponding to the first fault identifier from the configuration information as the initial handling method for the target fault. In other words, the configuration information is similar to a comprehensive database of the computer system, and the first fault identifier is the fault identifier corresponding to the target fault. Therefore, retrieving the handling method corresponding to the first fault identifier from the configuration information is the handling method for the target fault corresponding to the first fault identifier.

[0089] Of course, in other implementations, the target fault can be specifically analyzed to determine the handling method for the target fault, and the target fault can be handled using the handling method determined by the analysis.

[0090] In one implementation, the fault status of the target fault is also updated based on the processing result. By updating the fault status of the target fault, the completion of processing of the target fault is indicated, thus avoiding repeated processing of the target fault.

[0091] In one embodiment, information related to the handling of the target fault is recorded; wherein, the information related to handling includes at least one of the following: handling time, handling method, and handling result. By recording information related to the handling of the target fault, on the one hand, maintenance personnel can conduct regular analysis and summarization based on the information related to the fault handling, which can help identify common problems and potential risks in the computer system, thereby promoting continuous improvement and innovation of the computer system; on the other hand, the information related to the handling of the target fault can also be used to adjust the handling priority of other faults.

[0092] In one specific implementation, such as Figure 2 As shown, the lower-level machine also includes a fault handling module and a feedback and optimization module. Figure 2In the "Feedback and Optimization" section, after the fault handling module processes the target fault, it records the relevant information about the target fault and sends it to the feedback and optimization module. The feedback and optimization module also records the relevant information about the target fault.

[0093] In one embodiment, in response to a query request for the target fault issued by the host computer, fault-related information of the target fault is fed back to the host computer.

[0094] Please see Figure 5 , Figure 5 yes Figure 1 The flowchart shown is a schematic diagram of one embodiment of step S12. It should be noted that if substantially the same result is achieved, this embodiment does not necessarily follow the same pattern. Figure 5 The illustrated process sequence is limited. For example... Figure 5 As shown, the target data includes system parameters, which include at least two system state parameters. In this embodiment, these include:

[0095] Step S51: Obtain the weights corresponding to each system state parameter.

[0096] In this embodiment, the weights corresponding to each system state parameter are obtained. The weights of each system state parameter, including the system parameters, reflect their importance in the initial processing priority adjustment.

[0097] For example, the system state parameters involved in adjusting the initial processing priority of target fault A include CPU utilization, memory utilization, disk read / write speed, and data transmission speed. The weights corresponding to CPU utilization are 40%, memory utilization is 30%, disk read / write speed is 20%, and data transmission speed is 10%.

[0098] It should be noted that the system state parameters for adjusting the initial processing priority of different faults can be the same or different, and this is not limited here. Additionally, the weights of the system state parameters for adjusting the initial processing priority of different faults can be the same or different, and this is not limited here.

[0099] Step S52: Based on the weights corresponding to each system state parameter, perform weighted processing on each system state parameter to obtain the fault score of the target fault.

[0100] In this embodiment, the system state parameters are weighted according to their respective weights to obtain a fault score for the target fault. By weighting each system state parameter according to its corresponding weight, the impact of different system state parameters is fully considered, improving the accuracy and rationality of the determined target fault score. This, in turn, improves the accuracy and rationality of the subsequent target fault processing priority adjusted based on the fault score. Furthermore, the weights corresponding to the system state parameters reflect their importance in the initial processing priority adjustment. Therefore, by assigning higher weights to relatively more important system state parameters, their impact can be more fully considered, further improving the accuracy and rationality of the target fault score, and thus improving the accuracy and rationality of the adjusted processing priority.

[0101] In one implementation, before weighting each system state parameter based on its corresponding weight to obtain the fault score of the target fault, the system state parameters are normalized. This normalization ensures that system state parameters with different value ranges can be compared and calculated on the same scale.

[0102] In one specific implementation, linear normalization can be used to normalize the system state parameters. Of course, in other specific implementations, max-min normalization, decimal calibration normalization, etc., can also be used to normalize the system state parameters, and this is not limited here.

[0103] Step S53: Use the fault score to adjust the initial processing priority of the target fault.

[0104] In this embodiment, the initial processing priority of the target fault is adjusted using fault scores. The fault score of the target fault is obtained by weighting each system state parameter according to its corresponding weight, that is, it fully considers the influence of different system state parameters; therefore, the processing priority of the target fault adjusted based on the fault score is accurate and reasonable.

[0105] In one embodiment, the initial processing priority of the target fault is adjusted using the fault score, specifically: obtaining a reference processing priority corresponding to the fault score; in response to the reference processing priority being higher than the current processing priority of the target fault, adjusting the processing priority of the target fault to the reference processing priority, or increasing the processing priority of the target fault by a preset first adjustment step size; in response to the reference processing priority being lower than or equal to the current processing priority of the target fault, not adjusting the processing priority of the target fault, or decreasing the processing priority of the target fault by a preset second adjustment step size.

[0106] In other words, when the reference processing priority corresponding to the fault score of the target fault is higher than the current processing priority of the target fault, there are two adjustment methods: one is to adjust the processing priority of the target fault directly to the reference processing priority in one step, and the other is to gradually increase the processing priority of the target fault. When the reference processing priority corresponding to the fault score of the target fault is lower than the current processing priority of the target fault, there are two adjustment methods: one is not to adjust the processing priority of the target fault, and the other is to gradually decrease the processing priority of the target fault.

[0107] There are no restrictions on the first adjustment step size and the second adjustment step size. For example, both the first adjustment step size and the second adjustment step size are 1.

[0108] For example, taking the first adjustment step size as 1: the reference processing priority corresponding to the fault score of the target fault is "high", the current processing priority of the target fault is "low", and the processing priority of the target fault is increased to "medium".

[0109] For example, taking the second adjustment step size as 1: the reference processing priority corresponding to the fault score of the target fault is "low", the current processing priority of the target fault is "medium", and the processing priority of the target fault is lowered to "low".

[0110] In addition, no restrictions are placed on the reference processing priority corresponding to the fault score range. For example, the reference processing priority corresponding to a fault score in the range of 0.0-0.4 is low priority, the reference processing priority corresponding to a fault score in the range of 0.4-0.7 is medium priority, and the reference processing priority corresponding to a fault score in the range of 0.7-1.0 is high priority.

[0111] In one specific implementation, such as Figure 2 As shown, the lower-level machine also includes a system monitoring module ( Figure 2 The system status monitoring module collects system status parameters and feeds them back to the dynamic priority adjustment module. The dynamic priority adjustment module then executes steps S51-S53 based on the system status parameters collected by the system monitoring module.

[0112] Furthermore, the system monitoring module can periodically collect system status parameters, from which system status parameters related to the target fault can be obtained. Alternatively, the system monitoring module can also collect system status parameters when abnormalities are detected, and from which system status parameters related to the target fault can be obtained. For example, if memory usage is detected to be greater than 90%, the computer system's system status parameters are considered abnormal, and system status parameters are collected.

[0113] In one specific implementation, since a fault event for the target fault is generated in response to the detection of the target fault, the system state parameters at the time the fault event for the target fault is generated can be used as the system state parameters of the computer system when the target fault occurs, or the system state parameters after the fault event for the target fault is generated can be used as the system state parameters of the computer system after the target fault occurs.

[0114] In one specific implementation, the target data includes system status parameters of the computer system after the target fault occurs. These parameters can be system status parameters of the computer system over a certain period of time after the target fault occurs, or they can be system status parameters of the computer system at a specific point in time after the target fault occurs. No limitation is made here.

[0115] In one embodiment, the target data includes system parameters, which in turn include system performance parameters. Based on the target data, the initial processing priority of the target fault is adjusted. Specifically, in response to system performance parameters exceeding their range and the target fault being a preset associated fault of the system performance parameters, the initial processing priority of the target fault is increased. In other words, when the target fault is a preset associated fault of the system performance parameters, the initial processing priority of the target fault is increased, enabling subsequent target faults related to abnormal system parameters to be processed preferentially and promptly.

[0116] The system performance parameters are not limited in range; they can be set according to actual usage needs.

[0117] For example, the system performance parameter is the system load. When the system load exceeds the parameter range, it means that the system load does not meet the performance requirements. If the target fault is a preset associated fault of the system load, the initial processing priority of the target fault is increased. If the target fault is not a preset associated fault of the system load, the initial processing priority of the target fault is not adjusted.

[0118] It should be noted that when the system performance parameters exceed the parameter range and the target fault is a preset associated fault of the system performance parameters, the initial processing priority of the target fault can be increased by a certain adjustment step size so that the target fault can be processed first in the future.

[0119] In one specific implementation, such as Figure 2 As shown, the lower-level machine also includes a system monitoring module and a correlation analysis module. Figure 2 The system includes a "correlation analysis" module and a dynamic priority adjustment module. The system monitoring module collects system performance parameters and feeds them back to the correlation analysis module. The correlation analysis module determines whether the system performance parameters exceed the parameter range and whether the target fault is a preset correlated fault of the system performance parameters, thereby determining whether to increase the initial processing priority of the target fault. After determining to increase the initial processing priority of the target fault, it feeds back to the dynamic priority adjustment module, which then increases the initial processing priority of the target fault.

[0120] In one embodiment, the target data includes information affecting the fault; based on the target data, the processing priority of the target fault is adjusted, specifically: in response to the information affecting the fault meeting a preset condition, the initial processing priority of the target fault is increased, wherein, when the information affecting the fault includes historical fault processing results, the preset condition includes the number of historical fault processing results that are failures reaching a preset number threshold within a first historical time period; when the information affecting the fault includes the fault occurrence frequency, the preset condition includes the fault occurrence frequency being higher than a preset frequency threshold.

[0121] When the impacting fault information is based on historical fault handling results, a preset condition is that the number of failed historical fault handling results reaches a preset threshold within the first historical time period. In other words, if the target fault has been pending for a long time without resolution, its initial processing priority will be increased to ensure timely handling and resolution. Furthermore, even if automated processing fails again, the increased initial processing priority ensures priority and timely feedback to maintenance personnel for further processing, thus guaranteeing prompt resolution.

[0122] When the frequency of fault occurrence is the primary factor affecting fault information, the preset condition is that the frequency of fault occurrence exceeds a preset frequency threshold. In other words, if the target fault is one that occurs frequently within a short period of time, the initial processing priority of the target fault will be increased to ensure that the target fault is handled and resolved in a timely manner.

[0123] There are no restrictions on the size of the first historical time period, the preset quantity threshold, and the preset frequency threshold; these can be set according to actual usage needs.

[0124] It should be noted that when the fault information meets the preset conditions, the initial processing priority of the target fault can be increased. This can be done by increasing the initial processing priority of the target fault by a certain adjustment step, so that the target fault can be processed first in the future.

[0125] In one specific implementation, such as Figure 2 As shown, the lower-level machine also includes a feedback and optimization module and a dynamic priority adjustment module. The feedback and optimization module determines whether the number of failed historical faults in the first historical time period reaches a preset threshold, and / or determines whether the failure frequency of the target fault is higher than a preset frequency threshold, thereby determining whether to increase the initial processing priority of the target fault. After determining to increase the initial processing priority of the target fault, it feeds back to the dynamic priority adjustment module, which then increases the initial processing priority of the target fault.

[0126] In one embodiment, the target data includes fault association information. Based on the target data, the initial processing priority of the target fault is adjusted, specifically: in response to the fault association information indicating a preset correlation between the target fault and other faults, the initial processing priority of the target fault is increased. The preset correlation between the target fault and other faults indicates that the occurrence of the target fault will lead to the occurrence of other faults; that is, the target fault is the root cause fault that leads to the occurrence of other faults. Therefore, by increasing the initial processing priority of the target fault, it can be ensured that the target fault, as the root cause fault, is processed preferentially and promptly, avoiding the occurrence of other faults due to the failure to process the target fault, that is, avoiding the occurrence of a chain reaction, thereby reducing the impact on the computer system and ensuring the reliability and operational stability of the computer system.

[0127] In one specific implementation, such as Figure 6 As shown, Figure 6 This is a flowchart illustrating an embodiment of the method for determining fault association information of a target fault provided in this application. The target data includes fault association information, and the target fault has a preset correlation with other faults: the occurrence of the target fault will lead to the occurrence of other faults. The determination of the fault association information of the target fault specifically includes the following sub-steps:

[0128] Step S61: Generate several first transactions using historical fault data of the computer system.

[0129] In this embodiment, historical fault data of the computer system is used to generate several first transactions; wherein, the first transactions include several historical faults that occurred within a second historical time period, and the target fault exists in the several first transactions. That is to say, the historical fault data of all faults that occurred within the second historical time period form a transaction.

[0130] The duration of the second historical time period is not limited and can be set according to actual usage needs.

[0131] Step S62: Perform at least one statistical analysis on several first transactions to obtain at least one correlation characterization value between the target fault and other faults.

[0132] In this embodiment, at least one statistical method is performed on several first transactions to obtain at least one correlation degree characterization value between the target fault and other faults. Each correlation degree characterization value can characterize whether the occurrence of the target fault will lead to the occurrence of other faults.

[0133] In one implementation, such as Figure 7 As shown, Figure 7 yes Figure 6 The flowchart of step S62 shown is a schematic diagram of an embodiment. At least one statistical method is performed on several first transactions to obtain at least one correlation characterization value between the target fault and other faults. Specifically, it includes the following sub-steps:

[0134] Step S71: Select the first transaction containing the target fault from several first transactions, and use it as the second transaction.

[0135] In this embodiment, a first transaction containing the target fault is selected from a plurality of first transactions and used as a second transaction.

[0136] For example, there are first transactions A, B, C, and D. First transactions A and D contain the target fault, so first transactions A and D are considered as the second transaction.

[0137] Step S72: Calculate the proportion of other failures occurring in the second transaction, and use it as the confidence level between the target failure and other failures.

[0138] In this embodiment, the proportion of other faults occurring in the second transaction is statistically analyzed as the confidence level between the target fault and other faults.

[0139] For example, there are first transactions A, B, C, D, and E; first transactions A, C, and D contain the target fault, so first transactions A, C, and D are considered as the second transactions, and the number of second transactions is 3; since other faults only occur in first transaction A, the proportion of other faults occurring in the second transactions is 1 / 3, that is, the confidence level between the target fault and other faults is 1 / 3.

[0140] Step S73: Use the confidence level as a correlation characterization value; and / or, calculate the proportion of the combination of the target fault and other faults in a number of first transactions as the support between the target fault and other faults, and use the ratio between the confidence level and the support as a correlation characterization value.

[0141] In this embodiment, confidence level is used as a correlation characterization value; and / or, the proportion of combinations of the target fault and other faults occurring in several first transactions is statistically analyzed as the support between the target fault and other faults, and the ratio between confidence level and support is used as a correlation characterization value. Therefore, the correlation characterization value between the target fault and other faults can be flexibly set as confidence level and / or lift.

[0142] For example, if the confidence level between the target fault and other faults is 1 / 3, then 1 / 3 is used as a correlation coefficient.

[0143] For example, consider transactions A, B, C, D, and E. Transactions A and D contain both the target fault and other faults. Therefore, the proportion of the target fault and other fault combinations occurring in these transactions is 2 / 5, meaning the support between the target fault and other faults is 2 / 5. Since the confidence level between the target fault and other faults is 1 / 3, the ratio of confidence level to support, 5 / 6, is used as a correlation coefficient.

[0144] Step S63: Based on at least one correlation characterization value, determine whether there is a preset correlation between the target fault and other faults.

[0145] In this embodiment, based on the at least one correlation degree characterization value, it is determined whether there is a preset correlation between the target fault and other faults. Since the correlation degree characterization value characterizes whether the occurrence of the target fault will lead to the occurrence of other faults, it can be determined that there is a preset correlation between the target fault and other faults when the correlation degree characterization value characterizes that the occurrence of the target fault will lead to the occurrence of other faults.

[0146] In one embodiment, determining whether a predetermined correlation exists between a target fault and other faults based on at least one correlation degree characterization value specifically involves: in response to each correlation degree characterization value indicating that the occurrence of the target fault will lead to the occurrence of other faults, determining that a predetermined correlation exists between the target fault and other faults. That is, when each correlation degree characterization value indicates that the occurrence of the target fault will lead to the occurrence of other faults, it can be determined that a predetermined correlation exists between the target fault and other faults.

[0147] Of course, in other implementations, it is also possible to determine that there is a preset correlation between the target fault and other faults when one of the correlation characterization values ​​indicates that the occurrence of the target fault will lead to the occurrence of other faults.

[0148] In one specific embodiment, at least one correlation characterization value includes a confidence level. When the confidence level is greater than or equal to a confidence threshold, it indicates that the occurrence of the target fault will lead to the occurrence of other faults, and the probability of occurrence is equal to the magnitude of the confidence level. The magnitude of the confidence threshold is not limited; for example, the confidence threshold may be 0.6.

[0149] In one specific embodiment, at least one correlation characteristic value includes a lift degree. When the lift degree is greater than a lift threshold, it indicates that the occurrence of the target fault will lead to the occurrence of other faults. The lift threshold is not limited in size; for example, the lift threshold may be 1.

[0150] In one specific implementation, at least one correlation characterization value includes confidence and lift. When the confidence is greater than or equal to a confidence threshold and the lift is greater than a lift threshold, it indicates that the occurrence of the target fault will lead to the occurrence of other faults, and the probability of leading to the occurrence of other faults is the magnitude of the confidence.

[0151] In one implementation, the target data includes system status parameters, system performance parameters, fault impact information, and fault correlation information. This data is used to adjust the processing priority of the target fault in sequence. Each subsequent adjustment of the processing priority is based on the processing priority adjusted in the previous one.

[0152] Please see Figure 8 , Figure 8 This is a schematic diagram of an embodiment of the fault handling apparatus provided in this application. The fault handling apparatus 80 includes an acquisition module 81, an adjustment module 82, and a processing module 83. The acquisition module 81 is used to acquire target data related to the target fault when a target fault is detected, and to acquire the initial processing priority of the target fault from configuration information; the adjustment module 82 is used to adjust the initial processing priority of the target fault based on the target data; the processing module 83 is used to process the target fault according to the adjusted processing priority of the target fault.

[0153] The target data includes at least one of the following: system parameters of the computer system when or after the target fault occurs, fault information affecting the target fault, and fault association information. The fault association information indicates whether there is a pre-defined correlation between the target fault and other faults, and the fault information affecting the target fault can characterize the impact on the performance of the computer system.

[0154] The system parameters mentioned above include at least one of system status parameters and system performance parameters. The system status parameters include at least one of the following: CPU utilization, memory utilization, disk read / write speed, and data transmission speed. The system performance parameters include at least one of the following: system load, response time, and fault frequency. The fault-affecting information includes at least one of the following: historical fault handling results of the target fault and fault occurrence frequency. The target fault has a pre-defined correlation with other faults: the occurrence of the target fault will lead to the occurrence of other faults.

[0155] The target data includes system parameters, which include at least two system state parameters. The adjustment module 82 is used to adjust the initial processing priority of the target fault based on the target data, including: obtaining the weights corresponding to each system state parameter; performing weighted processing on each system state parameter based on the weights corresponding to each system state parameter to obtain the fault score of the target fault; and adjusting the initial processing priority of the target fault using the fault score.

[0156] The adjustment module 82 is used to perform weighted processing on each system state parameter based on the weights corresponding to each system state parameter to obtain the fault score of the target fault, including: normalizing each system state parameter; and / or, the adjustment module 82 is used to adjust the initial processing priority of the target fault using the fault score, including: obtaining the reference processing priority corresponding to the fault score; in response to the reference processing priority being higher than the current processing priority of the target fault, adjusting the processing priority of the target fault to the reference processing priority, or increasing the processing priority of the target fault by a preset first adjustment step size; and / or, in response to the reference processing priority being lower than or equal to the current processing priority of the target fault, decreasing the processing priority of the target fault by a preset second adjustment step size.

[0157] The target data mentioned above includes system parameters, which include system performance parameters. The adjustment module 82 is used to adjust the initial processing priority of the target fault based on the target data, including: in response to the system performance parameters exceeding the parameter range and the target fault being a preset associated fault of the system performance parameters, increasing the initial processing priority of the target fault.

[0158] The target data includes information affecting the fault; the adjustment module 82 is used to adjust the processing priority of the target fault based on the target data, including: in response to the information affecting the fault meeting a preset condition, increasing the initial processing priority of the target fault, wherein, when the information affecting the fault includes historical fault processing results, the preset condition includes the number of historical fault processing results that are failures reaching a preset number threshold within the first historical time period, and when the information affecting the fault includes the fault occurrence frequency, the preset condition includes the fault occurrence frequency being higher than a preset frequency threshold.

[0159] The target data includes fault association information; the adjustment module 82 is used to adjust the initial processing priority of the target fault based on the target data, including: in response to the fault association information indicating that the target fault has a preset correlation with other faults, increasing the initial processing priority of the target fault.

[0160] The target data includes fault association information, whereby the target fault exhibits a pre-defined correlation with other faults: the occurrence of the target fault leads to the occurrence of other faults. The steps for determining the fault association information of the target fault include: generating several first transactions using historical fault data of the computer system; wherein the first transactions include several historical faults occurring within a second historical time period, and the target fault exists in the several first transactions; performing at least one statistical analysis on the several first transactions to obtain at least one correlation degree characterization value between the target fault and other faults, each correlation degree characterization value characterizing whether the occurrence of the target fault leads to the occurrence of other faults; and determining whether there is a pre-defined correlation between the target fault and other faults based on at least one correlation degree characterization value.

[0161] The method involves performing at least one statistical analysis on several first transactions to obtain at least one correlation characterization value between the target fault and other faults, including: selecting a first transaction containing the target fault from the several first transactions as a second transaction; calculating the proportion of other faults occurring in the second transaction as the confidence level between the target fault and other faults; using the confidence level as a correlation characterization value; and / or calculating the proportion of combinations of the target fault and other faults occurring in the several first transactions as the support level between the target fault and other faults, and using the ratio between the confidence level and the support level as a correlation characterization value.

[0162] Specifically, determining whether there is a pre-defined correlation between the target fault and other faults based on at least one correlation degree characterization value includes: in response to each correlation degree characterization value indicating that the occurrence of the target fault will lead to the occurrence of other faults, determining that there is a pre-defined correlation between the target fault and other faults.

[0163] The acquisition module 81 is used to obtain the initial processing priority of the target fault from the configuration information, including: obtaining the first fault identifier of the target fault; querying the processing priority corresponding to the first fault identifier from the configuration information as the initial processing priority of the target fault. The configuration information contains mapping information of different faults of the computer system, and the mapping information includes the first fault identifier of the fault and the corresponding processing priority.

[0164] The processing module 83 is used to process the target fault before processing it according to the adjusted processing priority of the target fault, including: querying the processing method corresponding to the target fault from the configuration information; the processing module 83 is used to process the target fault according to the adjusted processing priority of the target fault, including: processing the target fault according to the adjusted processing priority of the target fault and processing it in the processing method.

[0165] The processing module 83 is used to process the target fault in a processing mode, including: in response to the processing mode being manual processing mode, feeding back the first fault-related information of the target fault to the host computer so that the host computer can display the first fault-related information through the first fault pop-up window; and in response to the processing mode being automatic processing mode, processing the target fault according to the automatic processing mode.

[0166] The processing module 83 is also used to respond to the failure of processing the target fault according to the automatic processing method and to feed back the second fault-related information of the target fault to the host computer.

[0167] The configuration information includes mapping information for different faults in the computer system. The mapping information includes the first fault identifier and the corresponding handling method. The processing module 83 is used to query the handling method corresponding to the target fault from the configuration information, including: obtaining the first fault identifier of the target fault; querying the handling method corresponding to the first fault identifier from the configuration information, and using it as the handling method corresponding to the target fault.

[0168] The first fault identifier consists of the device type identifier, board type identifier, module type identifier and second fault identifier of the target fault, and the second fault identifier is the identifier corresponding to the type of the target fault.

[0169] The processing module 83 is used to update the fault status of the target fault based on the processing result of the target fault; the processing module 83 is used to record the processing-related information of the target fault; wherein the processing-related information includes at least one of the following: processing time, processing method, and processing result; the acquisition module 81 is used to generate a fault event of the target fault in response to the detection of the target fault, and to feed the fault event back to the host computer so that the host computer can display the fault event of the target fault through a second fault pop-up window; in response to the query request for the target fault issued by the host computer, the acquisition module 81 sends the fault-related information of the target fault to the host computer; wherein the fault event includes at least one of the first fault identifier and mapping information of the target fault, and the mapping information includes at least one of the following: the first fault identifier of the target fault, the fault type of the target fault, the fault name of the target fault, the processing priority corresponding to the target fault, and the processing method corresponding to the target fault.

[0170] Please see Figure 9 , Figure 9 This is a schematic diagram of an embodiment of the electronic device provided in this application. The electronic device 90 includes a memory 91 and a processor 92 coupled to each other. The processor 92 is used to execute program instructions stored in the memory 91 to implement the steps of any of the above-described fault handling method embodiments. In a specific implementation scenario, the electronic device 90 may include, but is not limited to, a microcomputer or a server. In addition, the electronic device 90 may also include mobile devices such as laptops and tablets, which are not limited here.

[0171] Specifically, processor 92 controls itself and memory 91 to implement the steps of any of the above-described fault handling method embodiments. Processor 92 may also be referred to as a CPU (Central Processing Unit). Processor 92 may be an integrated circuit chip with signal processing capabilities. Processor 92 may also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor or any conventional processor. Furthermore, processor 92 may be implemented using integrated circuit chips.

[0172] Please see Figure 10 , Figure 10 This is a schematic diagram of an embodiment of the computer-readable storage medium provided in this application. The computer-readable storage medium 100 of this application embodiment stores program instructions 101. When executed, these program instructions 101 implement the methods provided by any embodiment of the fault handling method of this application and any non-conflicting combination thereof. The program instructions 101 can form a program file and be stored in the aforementioned computer-readable storage medium 100 in the form of a software product, so that a computer device (which may be a personal computer, server, or network device, etc.) can execute all or part of the steps of the methods of various embodiments of this application. The aforementioned computer-readable storage medium 100 includes various media capable of storing program code, such as a USB flash drive, mobile hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, or terminal devices such as computers, servers, mobile phones, and tablets.

[0173] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A failure handling method characterized by, The method is applied to a computer system and includes: If a target fault is detected, target data related to the target fault is obtained, and the initial processing priority of the target fault is obtained from the configuration information; Based on the target data, the initial processing priority of the target fault is adjusted; The target fault is processed according to the adjusted processing priority of the target fault.

2. The method according to claim 1, characterized in that, The target data includes at least one of the following: system parameters of the computer system when or after the target fault occurs, fault information affecting the target fault, and fault association information. The fault association information indicates whether the target fault has a preset correlation with other faults, and the fault information affecting the target fault can characterize the impact on the performance of the computer system.

3. The method of claim 2, wherein, The system parameters include at least one of system status parameters and system performance parameters. The system status parameters include at least one of the following: CPU utilization, memory utilization, disk read / write speed, and data transmission speed. The system performance parameters include at least one of the following: system load, response time, and failure frequency. The information affecting the fault includes at least one of the following: historical fault handling results of the target fault, and fault occurrence frequency; The target fault has a pre-defined correlation with other faults: the occurrence of the target fault will lead to the occurrence of the other faults.

4. The method according to any one of claims 2 to 3, characterized in that, The target data includes the system parameters, which include at least two system state parameters; the adjustment of the initial processing priority of the target fault based on the target data includes: Obtain the weights corresponding to each of the system state parameters; Based on the weights corresponding to each of the system state parameters, the system state parameters are weighted to obtain the fault score of the target fault. The initial processing priority of the target fault is adjusted using the fault score.

5. The method of claim 4, wherein, Before weighting the system state parameters based on their respective weights to obtain the fault score of the target fault, the method further includes: The system state parameters are normalized. And / or, adjusting the initial processing priority of the target fault using the fault score includes: Obtain the reference processing priority corresponding to the fault score; In response to the reference processing priority being higher than the current processing priority of the target fault, the processing priority of the target fault is adjusted to the reference processing priority, or the processing priority of the target fault is increased by a preset first adjustment step; and / or, In response to the reference processing priority being lower than or equal to the current processing priority of the target fault, the processing priority of the target fault is lowered according to a preset second adjustment step size.

6. The method according to any one of claims 2 to 5, characterized in that, The target data includes the system parameters, which include system performance parameters; adjusting the initial processing priority of the target fault based on the target data includes: In response to the system performance parameters exceeding the parameter range, and the target fault being a preset associated fault of the system performance parameters, the initial processing priority of the target fault is increased.

7. The method according to any one of claims 2 to 6, characterized in that, The target data includes the information affecting the fault; the adjustment of the initial processing priority of the target fault based on the target data includes: In response to the failure information meeting preset conditions, the initial processing priority of the target failure is increased. Wherein, when the failure information includes historical failure processing results, the preset conditions include the number of failed historical failure processing results reaching a preset number threshold within a first historical time period. When the failure information includes failure occurrence frequency, the preset conditions include the failure occurrence frequency being higher than a preset frequency threshold.

8. The method according to any one of claims 2 to 7, characterized in that, The target data includes the fault association information; the adjustment of the initial processing priority of the target fault based on the target data includes: In response to the fault association information indicating that the target fault has the preset association with the other faults, the initial processing priority of the target fault is increased.

9. The method according to any one of claims 2 to 8, characterized in that, The target data includes the fault association information, and the target fault has a preset correlation with other faults: the occurrence of the target fault will lead to the occurrence of the other faults; The steps for determining the fault association information of the target fault include: Using the historical fault data of the computer system, a plurality of first transactions are generated; wherein, the first transactions include a plurality of historical faults that occurred within a second historical time period, and the target fault exists in the plurality of first transactions; At least one statistical method is performed on the plurality of first transactions to obtain at least one correlation characterization value between the target fault and other faults, and each of the correlation characterization values ​​can characterize whether the occurrence of the target fault will lead to the occurrence of the other faults; Based on the at least one correlation characterization value, determine whether the target fault and the other faults have the preset correlation.

10. The method according to claim 9, characterized in that, The step of performing at least one statistical analysis on the plurality of first transactions to obtain at least one correlation characterization value between the target fault and other faults includes: From the plurality of first transactions, select the first transaction containing the target fault as the second transaction; The proportion of other failures occurring in the second transaction is used as the confidence level between the target failure and the other failures. The confidence level is used as a correlation characterization value; and / or, the proportion of the combination of the target fault and the other faults occurring in the plurality of first transactions is calculated as the support between the target fault and the other faults, and the ratio between the confidence level and the support is used as a correlation characterization value.

11. The method according to claim 9 or 10, characterized in that, Determining whether the target fault and the other faults have the preset correlation based on the at least one correlation characterization value includes: In response to each of the aforementioned correlation characterization values ​​indicating that the occurrence of the target fault will lead to the occurrence of the other faults, it is determined that the target fault and the other faults have the preset correlation.

12. The method according to any one of claims 1 to 11, characterized in that, The step of obtaining the initial processing priority of the target fault from the configuration information includes: Obtain the first fault identifier of the target fault; The processing priority corresponding to the first fault identifier is retrieved from the configuration information and used as the initial processing priority for the target fault. The configuration information contains mapping information for different faults of the computer system, and the mapping information includes the first fault identifier and the corresponding processing priority of the fault.

13. The method according to any one of claims 1 to 11, characterized in that, Before processing the target fault according to the adjusted processing priority, the method further includes: The processing method corresponding to the target fault can be obtained by querying the configuration information; The step of processing the target fault according to the adjusted processing priority of the target fault includes: The target fault is processed according to the processing priority adjusted based on the target fault, and the processing method is described above.

14. The method according to claim 13, characterized in that, The process of handling the target fault in the aforementioned manner includes: In response to the fact that the processing method is manual processing, the first fault-related information of the target fault is fed back to the host computer, so that the host computer displays the first fault-related information through the first fault pop-up window; In response to the fact that the processing method is an automatic processing method, the target fault is processed according to the automatic processing method.

15. The method according to claim 14, characterized in that, The method further includes: In response to the failure to process the target fault according to the automatic processing method, the system feeds back the second fault-related information of the target fault to the host computer.

16. The method according to any one of claims 13 to 15, characterized in that, The configuration information includes mapping information for different faults of the computer system, and the mapping information includes a first fault identifier and a corresponding handling method for the fault. The step of retrieving the processing method corresponding to the target fault from the configuration information includes: Obtain the first fault identifier of the target fault; The processing method corresponding to the first fault identifier is retrieved from the configuration information and used as the processing method corresponding to the target fault.

17. The method according to claim 12 or 16, characterized in that, The first fault identifier consists of the device type identifier, board type identifier, module type identifier and second fault identifier of the target fault, wherein the second fault identifier is the identifier corresponding to the type to which the target fault belongs.

18. The method according to any one of claims 1 to 17, characterized in that, The method further includes at least one of the following steps: Based on the processing results of the target fault, the fault status of the target fault is updated; Record the processing-related information of the target fault; wherein the processing-related information includes at least one of the following: processing time, processing method, and processing result; In response to detecting the target fault, a fault event for the target fault is generated and fed back to the host computer, so that the host computer displays the fault event for the target fault through a second fault pop-up window; in response to a query request for the target fault issued by the host computer, fault-related information of the target fault is fed back to the host computer; wherein, the fault event includes at least one of the first fault identifier and mapping information of the target fault, and the mapping information includes at least one of the following: the first fault identifier of the target fault, the fault type of the target fault, the fault name of the target fault, the processing priority corresponding to the target fault, and the processing method corresponding to the target fault.

19. A fault handling device, characterized in that, The device includes: The acquisition module is used to acquire target data related to the target fault when the target fault is detected, and to acquire the initial processing priority of the target fault from the configuration information; An adjustment module is used to adjust the initial processing priority of the target fault based on the target data; The processing module is used to process the target fault according to the adjusted processing priority of the target fault.

20. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory storing program instructions, and the processor executing the program instructions to implement the fault handling method as described in any one of claims 1-18.

21. A fault handling system, characterized in that, The fault handling system includes a host computer and a slave computer, wherein the slave computer is the electronic device described in claim 20.

22. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store program instructions that can be executed to implement the fault handling method as described in any one of claims 1-18.