Server fault processing method and device and electronic equipment

By configuring the fault drill scenario and quantifying the time difference score of the fault resolution steps, the problem of the server system's emergency operation and maintenance capabilities cannot be quantified, the emergency response efficiency is improved, and economic losses are reduced.

CN120448171APending Publication Date: 2025-08-08中国邮政储蓄银行股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510533072.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The emergency operation and maintenance capabilities of server systems in the prior art cannot be quantified, resulting in the inability to discover and deal with problems in the financial system, resulting in economic losses.

Method used

By obtaining the fault type of server failure, configuring the fault drill scenario, performing the fault drill and obtaining multiple fault resolution steps, determining the time difference for each fault resolution step, determining the solution score based on the time difference, and performing the fault resolution step to resolve the server failure when the total score is greater than the first preset threshold.

Benefits of technology

Quantitative assessment of emergency operation and maintenance capabilities has been achieved, emergency response efficiency has been improved, and economic losses have been reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448171A_ABST
    Figure CN120448171A_ABST
Patent Text Reader

Abstract

The invention provides a server fault processing method and device and electronic equipment, and the method comprises the steps: obtaining the fault type of a fault occurring in a server, and configuring a fault drilling scene according to the fault type; fault drilling is carried out in the fault drilling scene, a plurality of fault solving steps are obtained, the time difference of each fault solving step is determined, and the time difference represents the time consumed by each fault solving step; and determining a solution score of fault drilling according to each time difference, determining a total score according to the solution score, and executing a plurality of fault solving steps corresponding to a target fault type to solve the fault of the server at least under the condition that the total score is greater than a first preset threshold value and the server has the fault of the target fault type. Through the method and the device, the problem of economic loss caused by the fact that the emergency operation and maintenance capability of a server system cannot be quantified and the problems existing in a financial system cannot be found and processed in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of server failure processing, and in particular to a method, device, computer-readable storage medium, and electronic device for processing server failures. Background Art

[0002] In today's cloud-native financial landscape, the need to rapidly restore financial systems or equipment to normal operation in emergencies is paramount, minimizing business interruptions and ensuring business continuity and stability. Through timely and effective emergency response, financial losses and reputational damage caused by failures can be minimized or mitigated. Furthermore, user inconvenience and distress caused by system failures can be reduced, improving user satisfaction and loyalty. This places higher demands on the emergency operations capabilities of teams involved in such emergencies. Assessing emergency operations capabilities is a key tool for improving an organization's ability to respond to emergencies. Emergency operations include fault discovery and reporting, fault analysis and assessment, development of emergency response plans, execution of emergency response, fault recovery and verification, and process summary and improvement. Emergency drills simulate potential production failure scenarios, matching them to emergency operations processes. During this process, observation and acquisition of data on the team's emergency operations capabilities are conducted, quantitatively assessed, and continuous improvement measures are developed.

[0003] Existing techniques typically involve tabletop exercises, which discuss and rehearse the emergency decision-making and on-site response processes based on pre-set drill scenarios. These exercises focus on improving command, decision-making, and coordination, and help relevant personnel master the responsibilities and procedures stipulated in the emergency plan. Hands-on exercises, on the other hand, involve actual decision-making, actions, and operations, using pre-set emergency scenarios and subsequent developments to complete the actual emergency response process. While response time, response time, and recovery time can generally be quantitatively assessed, the completeness of emergency plans and their coordination and communication capabilities are often qualitatively assessed through organizational or expert review. Existing assessment methods rely on the subjectivity of expert review, often using ranges for emergency response time, making quantifiable scores difficult to assess. This leads to frequent manual evaluations after emergency drills, making them inefficient. Data-fusion drill assessment models lack mandatory and quantitative requirements for emergency response timelines, requiring specific implementation requirements for the plan. Furthermore, they lack a comprehensive emergency plan library for general emergency events. Therefore, a method is needed to improve the efficiency and quantification visibility of emergency operation and maintenance capability assessments. Summary of the Invention

[0004] The main purpose of this application is to provide a method, device, computer-readable storage medium and electronic device for handling server failures, so as to at least solve the problem in the prior art that the emergency operation and maintenance capabilities of the server system cannot be quantified, resulting in the inability to discover and handle problems in the financial system, thereby causing economic losses.

[0005] In order to achieve the above-mentioned purpose, according to one aspect of the present application, a method for handling server failures is provided, comprising: obtaining the failure type of the server failure, configuring a failure drill scenario according to the failure type, wherein the failure drill scenario includes at least the model of the server; performing a failure drill in the failure drill scenario and obtaining multiple failure resolution steps, determining the time difference of each failure resolution step, wherein the failure resolution step is a step for resolving the server failure, and the time difference represents the time consumed by each failure resolution step; determining a solution score for the failure drill according to each time difference, and determining a total score according to the solution score, at least when the total score is greater than a first preset threshold and a target failure type failure occurs on the server, executing multiple failure resolution steps corresponding to the target failure type to resolve the server failure, wherein the target failure type is one of the failure types.

[0006] Optionally, determining the time difference of each fault resolution step includes: obtaining the alarm reception time and the alarm issuance time, and calculating the difference between the alarm reception time and the alarm issuance time to obtain the alarm time difference, wherein the alarm issuance time indicates the time when the fault alarm is issued, and the alarm reception time indicates the time when the fault alarm is received; obtaining the fault location time, and calculating the difference between the fault location time and the alarm reception time to obtain the location time difference, wherein the fault location time indicates the time when the fault location is received; obtaining the fault handling time, and calculating the difference between the fault handling time and the fault location time to obtain the handling time difference, wherein the fault handling time indicates the time when the fault handling is completed.

[0007] Optionally, the solution score of the fault drill is determined according to each of the time differences, including: obtaining the total number of faults, and determining the number of alarms whose alarm time difference is less than a second preset threshold, calculating the ratio of the number of alarms to the total number of faults, and obtaining the alarm score, calculating the product of the ratio of the number of alarms to the total number of faults and the alarm score, and obtaining the alarm trigger score, wherein the alarm score represents the proportion of the fault alarm in the solution score; determining the number of positioning times whose positioning time difference is less than a third preset threshold, calculating the ratio of the number of positioning times to the total number of faults, and obtaining the positioning score, calculating the ratio of the number of positioning times to the total number of faults The product of the ratio of the number of faults to the number of located faults and the positioning score is used to obtain a positioning problem score, wherein the positioning score represents the proportion of fault positioning in the solution score; the number of located faults is obtained, the number of treatments whose treatment time difference is less than the fourth preset threshold is determined, the ratio of the number of treatments to the number of located faults is calculated, and the treatment score is obtained, the product of the ratio of the number of treatments to the number of located faults and the treatment score is calculated to obtain a treatment problem score, wherein the treatment score represents the proportion of fault treatment in the solution score; the sum of the alarm trigger score, the positioning problem score and the treatment problem score is calculated to obtain the solution score.

[0008] Optionally, determining a total score based on the solution score includes: determining a plan score, and determining a total answer score, wherein the plan score represents the score of the degree of compliance of the fault-solving steps with pre-set steps, and the total answer score is the score obtained by the user correctly answering the server fault problem; calculating the sum of the solution score, the plan score and the total answer score to obtain the total score.

[0009] Optionally, determining the contingency plan score includes: obtaining a preset processing solution, wherein the preset processing solution represents a plurality of historical fault resolution steps pre-set before the fault drill; calculating the ratio of the number of faults whose fault resolution steps are the same as the historical fault resolution steps to the total number of faults to obtain the contingency plan score.

[0010] Optionally, determining the total answer score includes: performing a fault drill in the fault drill scenario and obtaining multiple answer scores, calculating an average of the multiple answer scores, and obtaining the total answer score, wherein the answer score is the score of each user's answer.

[0011] Optionally, at least when the total score is greater than a first preset threshold and a target fault type occurs on the server, multiple fault resolution steps corresponding to the target fault type are executed to resolve the server fault, including: obtaining the total number of improvement measures and the total number of adopted improvement measures, wherein the improvement measures are improved measures for multiple fault resolution steps; calculating the ratio of the total number of adopted improvement measures to the total number of improvement measures to obtain an adoption score; when the adoption score is greater than a fifth preset threshold, updating the corresponding fault resolution step according to the adopted improvement measure, and executing each fault resolution step to resolve the server fault.

[0012] According to another aspect of the present application, a device for processing server failures is provided, comprising: a configuration unit, configured to obtain a failure type of a failure occurring in the server, and configure a failure drill scenario according to the failure type, wherein the failure drill scenario includes at least the model of the server; a determination unit, configured to perform a failure drill in the failure drill scenario and obtain multiple failure resolution steps, and determine a time difference for each of the failure resolution steps, wherein the failure resolution step is a step for resolving the failure of the server, and the time difference represents the time consumed by each of the failure resolution steps; an execution unit, configured to determine a solution score for the failure drill according to each of the time differences, and to determine a total score according to the solution score, and at least when the total score is greater than a first preset threshold and a failure of a target failure type occurs on the server, execute multiple failure resolution steps corresponding to the target failure type to resolve the failure of the server, wherein the target failure type is one of the failure types.

[0013] According to another aspect of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium includes a stored program, wherein when the program is run, the device where the computer-readable storage medium is located is controlled to execute any one of the server failure processing methods.

[0014] According to another aspect of the present application, an electronic device is provided, comprising: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include a method for executing any one of the server failure processing methods.

[0015] The technical solution of the present application is applied to obtain a fault type, configure a fault drill scenario based on the fault type, wherein the fault type indicates the type of server fault, and the fault drill scenario includes at least the server model; perform a fault drill in the fault drill scenario and obtain multiple fault resolution steps, determine the time difference of each fault resolution step, wherein the fault resolution step is a step for resolving the server fault, and the time difference indicates the time consumed by each fault resolution step; determine a solution score for the fault drill based on each time difference, and determine a total score based on the solution score; at least when the total score is greater than a first preset threshold and the server has a fault of the fault type, execute the multiple fault resolution steps to resolve the server fault. Compared with the prior art, in which the emergency operation and maintenance capabilities of the server system cannot be quantified, resulting in the inability to discover and address problems in the financial system, thereby causing economic losses, the present application scores the fault resolution steps through the above solution to quantify the problem-solving capabilities, and then determines whether to continue to use the fault resolution steps to handle the fault based on the quantified total score. If the total score is greater than the first preset threshold, the fault resolution steps are continued to be used to handle the fault, so as to timely discover and address problems in the server system. Therefore, it can solve the problem that the emergency operation and maintenance capabilities of the server system cannot be quantified, resulting in the inability to discover and deal with problems in the financial system, thereby causing economic losses, and achieve the effect of reducing economic losses. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings that constitute part of this application are used to provide a further understanding of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation on this application. In the drawings:

[0017] Figure 1 A hardware structure block diagram of a mobile terminal for executing a method for handling server failures provided in an embodiment of the present application is shown;

[0018] Figure 2 A schematic diagram showing a flow chart of a method for handling a server failure provided in an embodiment of the present application is shown;

[0019] Figure 3 A flowchart illustrating a specific method for handling server failures provided in an embodiment of the present application is shown;

[0020] Figure 4 A schematic diagram of a specific quantitative evaluation of server failure handling capability provided by an embodiment of the present application is shown;

[0021] Figure 5 A structural block diagram of a server failure processing device provided by an embodiment of the present application is shown.

[0022] The above drawings include the following reference numerals:

[0023] 102. Processor; 104. Memory; 106. Transmission device; 108. Input / output device. DETAILED DESCRIPTION

[0024] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0025] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application described here. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0027] For ease of description, some nouns or terms involved in the embodiments of the present application are explained below:

[0028] Emergency operation and maintenance: Emergency operation and maintenance refers to a series of emergency measures taken when an emergency occurs in a system or equipment in order to quickly troubleshoot the fault, restore the normal operation of the system or equipment, and try to recover or reduce the losses caused by the accident.

[0029] Emergency drills: Simulate emergencies or system failure scenarios and organize the operations team to handle them. Through these drills, we observe the team's response speed, collaboration, and recovery capabilities.

[0030] Capability Quantification Assessment: This quantitative evaluation assesses the emergency response, troubleshooting, recovery, and overall management capabilities of an organization's operations team when responding to emergencies or system failures. This assessment aims to understand the current state of an organization's emergency operations capabilities, identify existing problems and deficiencies, and develop improvement measures to enhance the organization's ability to respond to emergencies.

[0031] Red-Blue Confrontation: The "Blue Team" simulates emergencies or system failures, while the "Red Team" is responsible for responding and handling them. During the confrontation, both sides continuously improve the "Red Team's" ability to handle various emergencies.

[0032] As introduced in the background technology, the emergency operation and maintenance capabilities of the server system in the existing technology cannot be quantified, resulting in the inability to discover and handle problems in the financial system, thereby causing economic losses. In order to solve the problem of the inability to quantify emergency operation and maintenance capabilities, the embodiments of the present application provide a server failure processing method, device, computer-readable storage medium and electronic device.

[0033] The technical solutions in the embodiments of the present invention will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present invention.

[0034] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 FIG. 1 is a hardware structure diagram of a mobile terminal of a method for handling a server failure according to an embodiment of the present invention. Figure 1 As shown, the mobile terminal may include one or more ( Figure 1 Only one is shown) a processor 102 (the processor 102 may include but is not limited to a microprocessor MCU or a programmable logic device FPGA and other processing devices) and a memory 104 for storing data, wherein the mobile terminal may also include a transmission device 106 and an input and output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the mobile terminal. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0035] The memory 104 can be used to store computer programs, such as software programs and modules of application software, such as the computer program corresponding to the server failure handling method in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implementing the above-mentioned method. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories can be connected to the mobile terminal via a network. Examples of the above-mentioned networks include but are not limited to the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by the mobile terminal's communication provider. In one example, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In one example, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0036] In this embodiment, a method for handling server failures running on a mobile terminal, a computer terminal, or a similar computing device is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0037] Figure 2 FIG. 1 is a flow chart of a method for handling server failure according to an embodiment of the present application. Figure 2 As shown, the method includes the following steps:

[0038] Step S201: Obtain the fault type of the server fault, and configure a fault drill scenario according to the fault type, wherein the fault drill scenario at least includes the model of the server;

[0039] Specifically, identify the types of failures that the server may encounter, including but not limited to hardware failures, software failures, network failures, etc. Information on the type of failure is crucial for the targeted nature of the drill. After determining the type of failure, the system of the present invention is used to configure the corresponding failure drill scenario. The scenario configuration needs to take into account the specific details of the failure, such as the severity of the failure, the possible scope of impact, and the server models involved. This is to ensure that the drill scenario can truly reflect the possible emergency situations, so that the emergency response team can experience a failure environment close to the real one during the drill.

[0040] Step S202: Perform a fault drill in the fault drill scenario and obtain multiple fault resolution steps, and determine a time difference between each of the fault resolution steps, wherein the fault resolution step is a step for resolving the server fault, and the time difference represents the time consumed by each of the fault resolution steps.

[0041] Specifically, during a fault drill, the emergency response team conducts a drill according to a pre-set troubleshooting process. During the drill, the system records the execution time of each key step, including the time of receiving the alarm, locating the problem, and handling the problem. The time difference is calculated as the difference between the start and end time of each troubleshooting step, which directly reflects the team's response speed and processing efficiency at each step.

[0042] Step S203: Determine a solution score for the fault drill based on each of the above time differences, and determine a total score based on the above solution score. At least when the above total score is greater than a first preset threshold and a target fault type fault occurs on the server, execute multiple fault resolution steps corresponding to the target fault type to resolve the server fault, wherein the target fault type is one of the above fault types.

[0043] Specifically, based on the time difference obtained in the drill, the system uses preset scoring rules and algorithms to calculate the solution score of the fault drill. The scoring rules may include specific requirements for response speed, positioning accuracy, handling efficiency, etc. The solution score is combined with other evaluation indicators, such as the hit rate of the plan, the score of the answer, etc., to comprehensively calculate the total score. The total score is a value that reflects the overall emergency response capability of the emergency response team, where the first preset threshold is used as the standard for the team to reach a qualified emergency response level. When the total score of the emergency response team in the drill exceeds the first preset threshold, and in actual operation and maintenance, the server does have the same or similar target fault type as in the drill, the system recommends executing the fault resolution steps corresponding to the fault type and having a higher score in the drill. Based on the feedback mechanism of the scoring results, by selecting the handling plan that performed well in the drill, it can guide the actual fault handling and improve the efficiency and success rate of fault resolution.

[0044] Through this embodiment, a fault type is obtained, and a fault drill scenario is configured based on the fault type. The fault type represents the type of server fault, and the fault drill scenario includes at least the server model. A fault drill is performed in the fault drill scenario and multiple fault resolution steps are obtained. The time difference of each fault resolution step is determined. The fault resolution step is a step for resolving the server fault, and the time difference represents the time consumed by each fault resolution step. A solution score for the fault drill is determined based on each time difference, and a total score is determined based on the solution score. At least when the total score is greater than a first preset threshold and the server has a fault of the fault type, the multiple fault resolution steps are executed to resolve the server fault. Compared with the prior art, in which the emergency operation and maintenance capabilities of the server system cannot be quantified, resulting in the inability to discover and address problems in the financial system, thereby causing economic losses, the present application scores the fault resolution steps through the above scheme to quantify the problem-solving capabilities. The quantified total score is then used to determine whether to continue using the fault resolution steps to address the fault. If the total score is greater than the first preset threshold, the fault resolution steps are continued to address the fault, thereby promptly discovering and addressing problems in the server system. Therefore, it can solve the problem that the emergency operation and maintenance capabilities of the server system cannot be quantified, resulting in the inability to discover and deal with problems in the financial system, thereby causing economic losses, and achieve the effect of reducing economic losses.

[0045] In the specific implementation process, the above-mentioned step S202 determines that the time difference of each of the above-mentioned fault resolution steps can be achieved by the following steps: Step S2021: Obtain the alarm reception time and the alarm issuance time, and calculate the difference between the above-mentioned alarm reception time and the above-mentioned alarm issuance time to obtain the alarm time difference, wherein the above-mentioned alarm issuance time indicates the time when the fault alarm is issued, and the above-mentioned alarm reception time indicates the time when the above-mentioned fault alarm is received; Step S2022: Obtain the fault location time, and calculate the difference between the above-mentioned fault location time and the above-mentioned alarm reception time to obtain the location time difference, wherein the above-mentioned fault location time indicates the time when the fault location is received; Step S2023: Obtain the fault handling time, and calculate the difference between the above-mentioned fault handling time and the above-mentioned fault location time to obtain the handling time difference, wherein the above-mentioned fault handling time indicates the time when the above-mentioned fault handling is completed. This method calculates the time difference of each step through the above-mentioned steps to quantitatively evaluate the efficiency and capabilities of the emergency response team.

[0046] Specifically, the alarm time difference is as follows: Obtain the alarm issuance time: In a simulated fault scenario, the system (blue team) or the drill platform automatically triggers a fault alarm, and the exact time of alarm triggering is recorded. Obtain the alarm reception time: Record the time when the emergency response team (red team) receives the fault alarm information. This is usually achieved through the emergency drill platform or system log tracking. Calculate the alarm time difference: Use subtraction to calculate the time interval between the alarm issuance time and the alarm reception time to obtain the alarm time difference. This time difference reflects the team's response speed to the fault alarm and is a key indicator for evaluating emergency response capabilities. Localization time difference: Obtain the fault location time: Record the time when the emergency response team determines the specific location or cause of the fault. This step usually requires high technical skills and experience. Calculate the location time difference: Use subtraction to calculate the time interval between the fault location time and the alarm reception time to obtain the location time difference. This time difference measures the efficiency of the team's fault analysis and location after receiving the alarm. Resolve time difference: Obtain the fault resolve time: Record the time when the fault is completely resolved and the system returns to normal operation. Calculate the time difference between the fault resolution and fault location by subtracting the time difference between the fault resolution and fault location. This time difference reflects the team's efficiency in taking action and resolving the problem after locating the fault.

[0047] Table 1 Server system emergency drill record

[0048]

[0049] Table 1 shows a server system emergency drill record table, including columns such as drill date, drill personnel, and system version. The following information is filled in according to the serial number: alarm step: alarm receipt time, alarm information, and screenshot; positioning step: positioning end time, positioning problem, and screenshot; handling step: emergency handling end time, emergency handling description, and screenshot. Then, the alarm time difference, positioning time difference, and handling time difference are calculated.

[0050] In some optional implementations, step S2023 determines the solution score of the above-mentioned fault drill based on each of the above-mentioned time differences, which can be achieved by the following steps: step S2024: obtaining the total number of faults, and determining the number of alarms whose alarm time difference is less than the second preset threshold, calculating the ratio of the above-mentioned number of alarms to the above-mentioned total number of faults, and obtaining the alarm score, calculating the product of the ratio of the above-mentioned number of alarms to the above-mentioned total number of faults and the above-mentioned alarm score, and obtaining the above-mentioned alarm trigger score, wherein the above-mentioned alarm score represents the proportion of the fault alarm in the above-mentioned solution score; step S2025: determining the number of positioning times whose positioning time difference is less than the third preset threshold, calculating the ratio of the above-mentioned number of positioning times to the above-mentioned total number of faults, and obtaining the positioning score, and calculating the above-mentioned positioning score. The product of the ratio of the number of bits to the total number of faults and the positioning score is obtained to obtain the positioning problem score, wherein the positioning score represents the proportion of fault positioning in the solution score; step S2026: obtain the number of located faults, determine the number of treatments whose treatment time difference is less than the fourth preset threshold, calculate the ratio of the treatment number to the number of located faults, and obtain the treatment score, calculate the product of the ratio of the treatment number to the number of located faults and the treatment score to obtain the treatment problem score, wherein the treatment score represents the proportion of fault treatment in the solution score; step S2027: calculate the sum of the alarm trigger score, the positioning problem score and the treatment problem score to obtain the solution score. Through the above steps, this method clarifies the efficiency requirements of the three stages of alarm response, fault location and fault treatment, and calculates the scores in combination with the scoring rules, making the evaluation of emergency response capabilities more objective and quantitative.

[0051] Specifically, the calculation of the alarm trigger score: First, obtain the total number of failures in the entire emergency drill process, which is a quantitative indicator that represents all possible failure conditions set in the drill. Then, determine the number of alarms whose alarm time difference is less than the second preset threshold, that is, those situations where the response speed is faster than the preset standard after the failure occurs. Calculate the ratio of the number of alarms to the total number of failures. This step evaluates the team's ability to respond quickly to failures. Next, multiply it by the alarm score. The alarm score here represents the proportion of fault alarms in the entire emergency drill scoring system. For example, the alarm score is 10 points. Finally, get the alarm trigger score, which reflects the efficiency and performance of the team in responding to alarms.

[0052] Calculating the Location Problem Score: In this step, we first obtain the total number of faults in the drill. Next, we find the number of locations where the location time difference is less than the third preset threshold, that is, the number of cases in which the team successfully located the fault within the specified time. By calculating the ratio of the number of locations to the total number of faults, we evaluate the team's efficiency in fault location. This is then multiplied by the location score, which represents the importance of fault location in the overall assessment. Finally, we obtain the location problem score, which measures the team's fault location capabilities and response speed. For example, a location score of 20 points is sufficient.

[0053] Calculating the Problem Handling Score: First, obtain the number of faults that need to be located during the drill, i.e., the number of located faults. Next, find the number of faults whose resolution time difference is less than the fourth preset threshold. This indicates that the team was able to complete fault handling within the preset time. By calculating the ratio of the number of resolved faults to the number of located faults, the team's efficiency and ability in fault handling are assessed. This is then multiplied by the resolution score, which reflects the weight of fault handling in the overall evaluation system. Finally, the Problem Handling Score is obtained, which reflects the team's performance during the fault handling phase. For example, a resolution score of 30 points is used. Determining the Final Solution Score: The three scores mentioned above—the alarm trigger score, the problem location score, and the problem handling score—are added together to obtain the solution score for the entire fault drill. This total score intuitively demonstrates the emergency response team's comprehensive performance in the three key aspects of rapid response, accurate location, and efficient resolution.

[0054] In some optional implementations, the above-mentioned step S203 determines the total score based on the above-mentioned solution score, which can be achieved by the following steps: Step S2031: Determine the plan score and determine the total answer score, wherein the above-mentioned plan score represents the score of the degree of compliance of the above-mentioned troubleshooting steps with the pre-set steps, and the above-mentioned total answer score is the score obtained by the user correctly answering the above-mentioned server fault problem; Step S2032: Calculate the sum of the above-mentioned solution score, the above-mentioned plan score and the above-mentioned total answer score to obtain the above-mentioned total score. This method comprehensively considers multiple dimensions of emergency drills through the calculation of the total score, not only evaluates the actual operational capabilities of emergency handling, but also considers the team's implementation of the plan and the members' mastery of theoretical knowledge, providing a comprehensive and in-depth assessment of emergency response capabilities.

[0055] Specifically, the plan score is determined by assessing the degree to which the troubleshooting steps align with the pre-defined emergency response plan. During emergency drills, the team's troubleshooting methods are compared with the pre-defined plan. The plan score is designed to verify whether the team adhered to standard emergency procedures during the drill and the practicality of the plan. A high plan score indicates that the team was able to skillfully and accurately execute the plan, reflecting the plan's feasibility and the team's execution capabilities. The overall score is the score of the user's knowledge points tested during the emergency drill, including identification of fault types, mastery of emergency response procedures, and understanding of troubleshooting strategies. This step tests team members' theoretical knowledge and practical skills to ensure the team's technical ability to handle various faults. A high overall score indicates that team members have a solid grasp of theoretical knowledge of emergency response, providing a solid theoretical foundation for actual troubleshooting. The overall score is calculated by adding the solution score, the plan score, and the overall score. The solution score reflects the team's efficiency and capabilities in actual troubleshooting drills, the plan score verifies the effectiveness of the plan's execution, and the overall score reflects the team members' theoretical knowledge. By combining these three, the overall score comprehensively reflects the emergency response team's performance in all aspects of emergency handling, including rapid response, accurate plan execution, and in-depth theoretical knowledge. Table 2 shows a quantitative scoring table. The total score is the calculation method for this overall score, where P0 ≥ P1 ≥ P2 ≥ P3 ≥ P4, P4 ≤ P5, and P5 ≥ P6.

[0056] Table 2 Quantitative scoring table

[0057]

[0058]

[0059] In some optional implementations, step S2031 of determining the emergency plan score can be achieved by: obtaining a preset handling solution, where the preset handling solution represents a plurality of historical fault resolution steps pre-set prior to the fault drill; and calculating the ratio of the number of faults for which the fault resolution steps are identical to the historical fault resolution steps to the total number of faults to obtain the emergency plan score. This method verifies the practicality and implementation effectiveness of the emergency plan through the above steps, promotes the standardization and consistency of the team's emergency response capabilities, and thus provides an effective assessment and improvement tool for improving overall emergency operation and maintenance capabilities.

[0060] Specifically, before the emergency drill begins, the system or team will pre-define a set of solutions (i.e., historical troubleshooting steps) based on historical fault data and expert experience. These solutions cover a variety of possible fault types and corresponding handling processes. These pre-determined solutions serve as guidelines for the team to follow when facing specific faults. They aim to improve the efficiency and success rate of fault handling through standardized and regularized emergency response processes. During the drill, the team will execute a series of troubleshooting steps based on the fault they encounter. The system records these actual troubleshooting steps and compares them with the historical troubleshooting steps in the pre-determined solutions. Specifically, the system counts the number of troubleshooting steps executed during the drill that exactly match the pre-determined solutions, then divides this number by the total number of faults encountered during the drill to calculate a ratio. This ratio reflects the degree to which the troubleshooting steps executed by the team during the drill align with the requirements of the emergency plan. Once this ratio is calculated, the system converts it into a plan score, which directly reflects the degree to which the team's emergency response process aligns with the pre-determined solutions. The higher the score, the more the emergency measures taken by the team during the drill are in line with historical experience and plans developed by experts, demonstrating the team's good grasp of the plan and its ability to execute it.

[0061] In some optional implementations, step S2031 of determining the total score for the answers can be accomplished by performing a fault drill in the aforementioned fault drill scenario and obtaining multiple scores for the answers, calculating the average of the multiple scores to obtain the total score, where the scores represent the scores of each user's answers. This method, based on user scores, can identify areas where team members have knowledge gaps or misunderstandings, providing guidance for subsequent training and knowledge supplementation.

[0062] Specifically, the emergency drill platform simulates a series of scenarios related to possible server failures. These scenarios include hardware failures, such as hard drive and memory failures, as well as software failures, such as operating system anomalies, application crashes, and network outages. In these scenarios, users are asked to identify the failure type, analyze the cause, and propose solutions, or answer questions about troubleshooting procedures, standards, and best practices. During the drill, each user independently completes a series of theoretical knowledge tests, which may include multiple-choice, fill-in-the-blank, and short-answer questions. The user's answers are immediately scored by the system, based on factors such as the accuracy of the answer, the rationality of the solution, and familiarity with emergency response procedures. Each user receives a score based on their performance. The system collects all user scores and then averages them to create an overall score for the entire team. This average calculation is simple and intuitive, providing a comprehensive reflection of the team members' overall theoretical knowledge and mastery of troubleshooting knowledge. In addition to existing indicators (such as response time, location efficiency, handling speed, plan compliance, and continuous improvement capabilities), additional evaluation indicators such as resource utilization, user feedback, and economic costs can be added. For example, consider whether emergency resources are reasonably allocated and utilized during the emergency response process, user satisfaction, and the economic impact of emergency handling on the enterprise. Develop corresponding indicator collection modules to collect data related to the newly added evaluation dimensions, such as resource consumption records, user satisfaction survey results, and economic cost estimates. Through data analysis and mining techniques, these new indicators can be integrated into the existing quantitative evaluation model to achieve a more comprehensive performance evaluation.

[0063] In some optional implementations, the above-mentioned step S203 can also be implemented through the following steps: step S2033: obtaining the total number of improvement measures and the total number of the above-mentioned improvement measures that have been adopted, wherein the above-mentioned improvement measures are measures for improving multiple of the above-mentioned troubleshooting steps; step S2034: calculating the ratio of the total number of the above-mentioned improvement measures that have been adopted to the total number of the above-mentioned improvement measures to obtain an adoption score; step S2035: when the above-mentioned adoption score is greater than the fifth preset threshold, updating the corresponding above-mentioned troubleshooting steps according to the above-mentioned improvement measures that have been adopted, and executing each of the above-mentioned troubleshooting steps to solve the above-mentioned server fault. This method promotes the regular optimization and updating of emergency plans and troubleshooting steps by associating the adoption score with improvement measures, thereby improving the applicability and efficiency of the plans.

[0064] Specifically, the total number of improvement measures and the total number of adopted improvement measures are captured. After each emergency drill, the team reviews the drill, summarizes lessons learned, and proposes improvement measures. These improvement measures may include optimizing troubleshooting procedures, adjusting emergency response plans, and improving resource allocation. The system records the total number of proposed improvement measures and the number of those subsequently adopted and implemented. This step provides the basic data for calculating the adoption score. By calculating the ratio of the total number of adopted improvement measures to the total number of improvement measures, the system generates an adoption score. The adoption score reflects the team's attention to emergency drill feedback and the implementation rate of improvement measures. A high adoption score indicates that the team effectively learned from the drill and converted feedback into concrete improvement actions. When the adoption score exceeds the fifth preset threshold, it indicates a high adoption rate and that the team actively applies post-drill feedback to optimize troubleshooting procedures. Based on the adopted improvement measures, the system automatically or assists the team in updating troubleshooting procedures, ensuring that plans and emergency response strategies reflect the latest operational knowledge and best practices. Once the troubleshooting steps have been updated, the team will execute these optimized steps to resolve actual server failures, particularly those matching the target failure types in the drill. The updated steps are more efficient and precise, helping the team perform better in actual troubleshooting, reducing recovery time and improving system stability. Table 3 shows a statistical table for continuous improvement measures, used to record improvement measures.

[0065] Table 3 Statistics of continuous improvement measures

[0066]

[0067] In order to enable those skilled in the art to more clearly understand the technical solution of the present application, the implementation process of the server failure processing method of the present application will be described in detail below with reference to specific embodiments.

[0068] This embodiment relates to a specific method for handling server failures, such as Figure 3 As shown, the following steps are included:

[0069] Step S1: The Blue Team obtains information such as the emergency plan and operation and maintenance test papers, and creates a new drill plan;

[0070] Step S2: The referee can view the emergency plan, the drill plan, and the operation and maintenance test paper;

[0071] Step S3: The Red Team executes the drill plan and conducts operation and maintenance questions;

[0072] Step S4: The Red Team uploads the emergency record and saves it;

[0073] Step S5: The blue team conducts quantitative evaluation and scoring;

[0074] Step S6: The referee gives the players some improvement measures for the training, and the red team and the blue team each decide whether to adopt or not;

[0075] Step S7: Generate a drill report;

[0076] Step S8: End.

[0077] This embodiment relates to a specific quantitative evaluation diagram of the processing capacity of server failures, such as Figure 4 As shown, the assessment is conducted along six dimensions: process mastery, rapid response, accurate positioning, timely disposal, complete plans, and continuous improvement. Among them, operational level: the capabilities at this level focus on process mastery and rapid response, indicating that the capabilities of the emergency operation and maintenance team (including the server's emergency operation and maintenance equipment and personnel) in positioning and disposal and other process specifications need to be improved. Standardization level: Based on the above operational level, the accurate positioning and emergency disposal capabilities can meet expectations, but the plans are not complete enough and continuous improvement is poor, indicating that the emergency operation and maintenance team's collaborative foresight and adaptability are not growing enough. Proficient level: All six dimensions have good performance, but have not yet met expectations. The overall emergency operation and maintenance capabilities need to be improved as a whole. Excellent level: The emergency operation and maintenance capability assessment is close to expectations, with excellent performance in all dimensions, indicating that the operation and maintenance team has professional plans and standardized operations, can maintain the vitality of continuous innovation, and can calmly deal with emergency operation and maintenance incidents.

[0078] The embodiment of the present application also provides a device for processing a server failure. It should be noted that the device for processing a server failure in the embodiment of the present application can be used to execute the method for processing a server failure provided in the embodiment of the present application. The device is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware for a predetermined function. Although the device described in the following embodiments is preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceived.

[0079] The following introduces the server failure processing device provided in the embodiment of the present application.

[0080] Figure 5 Schematic diagram of a server failure processing device according to an embodiment of the present application. Figure 5 As shown, the device includes:

[0081] The configuration unit 10 is configured to obtain a fault type of a fault occurring on the server and configure a fault drill scenario according to the fault type, wherein the fault drill scenario at least includes a model of the server;

[0082] Specifically, identify the types of failures that the server may encounter, including but not limited to hardware failures, software failures, network failures, etc. Information on the type of failure is crucial for the targeted nature of the drill. After determining the type of failure, the system of the present invention is used to configure the corresponding failure drill scenario. The scenario configuration needs to take into account the specific details of the failure, such as the severity of the failure, the possible scope of impact, and the server models involved. This is to ensure that the drill scenario can truly reflect the possible emergency situations, so that the emergency response team can experience a failure environment close to the real one during the drill.

[0083] a determining unit 20 configured to perform a fault drill in the fault drill scenario and obtain multiple fault resolution steps, and determine a time difference between each of the fault resolution steps, wherein the fault resolution step is a step for resolving the server fault, and the time difference represents the time consumed by each of the fault resolution steps;

[0084] Specifically, during a fault drill, the emergency response team conducts a drill according to a pre-set troubleshooting process. During the drill, the system records the execution time of each key step, including the time of receiving the alarm, locating the problem, and handling the problem. The time difference is calculated as the difference between the start and end time of each troubleshooting step, which directly reflects the team's response speed and processing efficiency at each step.

[0085] The execution unit 30 is used to determine the solution score of the above-mentioned fault drill based on each of the above-mentioned time differences, and determine the total score based on the above-mentioned solution score, at least when the above-mentioned total score is greater than a first preset threshold and a fault of the target fault type occurs on the server, execute the multiple above-mentioned fault resolution steps corresponding to the above-mentioned target fault type to resolve the fault of the above-mentioned server, wherein the above-mentioned target fault type is one of the above-mentioned fault types.

[0086] Specifically, based on the time difference obtained in the drill, the system uses preset scoring rules and algorithms to calculate the solution score of the fault drill. The scoring rules may include specific requirements for response speed, positioning accuracy, handling efficiency, etc. The solution score is combined with other evaluation indicators, such as the hit rate of the plan, the score of the answer, etc., to comprehensively calculate the total score. The total score is a value that reflects the overall emergency response capability of the emergency response team, where the first preset threshold is used as the standard for the team to reach a qualified emergency response level. When the total score of the emergency response team in the drill exceeds the first preset threshold, and in actual operation and maintenance, the server does have the same or similar target fault type as in the drill, the system recommends executing the fault resolution steps corresponding to the fault type and having a higher score in the drill. Based on the feedback mechanism of the scoring results, by selecting the handling plan that performed well in the drill, it can guide the actual fault handling and improve the efficiency and success rate of fault resolution.

[0087] Through this embodiment, a fault type is obtained, and a fault drill scenario is configured based on the fault type. The fault type represents the type of server fault, and the fault drill scenario includes at least the server model. A fault drill is performed in the fault drill scenario and multiple fault resolution steps are obtained. The time difference of each fault resolution step is determined. The fault resolution step is a step for resolving the server fault, and the time difference represents the time consumed by each fault resolution step. A solution score for the fault drill is determined based on each time difference, and a total score is determined based on the solution score. At least when the total score is greater than a first preset threshold and the server has a fault of the fault type, the multiple fault resolution steps are executed to resolve the server fault. Compared with the prior art, in which the emergency operation and maintenance capabilities of the server system cannot be quantified, resulting in the inability to discover and address problems in the financial system, thereby causing economic losses, the present application scores the fault resolution steps through the above scheme to quantify the problem-solving capabilities. The quantified total score is then used to determine whether to continue using the fault resolution steps to address the fault. If the total score is greater than the first preset threshold, the fault resolution steps are continued to address the fault, thereby promptly discovering and addressing problems in the server system. Therefore, it can solve the problem that the emergency operation and maintenance capabilities of the server system cannot be quantified, resulting in the inability to discover and deal with problems in the financial system, thereby causing economic losses, and achieve the effect of reducing economic losses.

[0088] In the specific implementation process, the above-mentioned determination unit includes a first calculation module, a second calculation module and a third calculation module. The first calculation module is used to obtain the alarm reception time and the alarm issuance time, and calculate the difference between the above-mentioned alarm reception time and the above-mentioned alarm issuance time to obtain the alarm time difference, wherein the above-mentioned alarm issuance time indicates the time when the fault alarm is issued, and the above-mentioned alarm reception time indicates the time when the above-mentioned fault alarm is received; the second calculation module is used to obtain the fault location time, and calculate the difference between the above-mentioned fault location time and the above-mentioned alarm reception time to obtain the location time difference, wherein the above-mentioned fault location time indicates the time when the fault location is received; the third calculation module is used to obtain the fault handling time, and calculate the difference between the above-mentioned fault handling time and the above-mentioned fault location time to obtain the handling time difference, wherein the above-mentioned fault handling time indicates the time when the above-mentioned fault handling is completed. The device calculates the time difference of each step through the above-mentioned steps to quantitatively evaluate the efficiency and capabilities of the emergency response team.

[0089] Specifically, the alarm time difference is as follows: Obtain the alarm issuance time: In a simulated fault scenario, the system (blue team) or the drill platform automatically triggers a fault alarm, and the exact time of alarm triggering is recorded. Obtain the alarm reception time: Record the time when the emergency response team (red team) receives the fault alarm information. This is usually achieved through the emergency drill platform or system log tracking. Calculate the alarm time difference: Use subtraction to calculate the time interval between the alarm issuance time and the alarm reception time to obtain the alarm time difference. This time difference reflects the team's response speed to the fault alarm and is a key indicator for evaluating emergency response capabilities. Localization time difference: Obtain the fault location time: Record the time when the emergency response team determines the specific location or cause of the fault. This step usually requires high technical skills and experience. Calculate the location time difference: Use subtraction to calculate the time interval between the fault location time and the alarm reception time to obtain the location time difference. This time difference measures the efficiency of the team's fault analysis and location after receiving the alarm. Resolve time difference: Obtain the fault resolve time: Record the time when the fault is completely resolved and the system returns to normal operation. Calculating the resolution time difference: Using subtraction, calculate the time interval between the fault resolution moment and the fault location moment to obtain the resolution time difference. The resolution time difference reflects the efficiency of the team in taking measures and resolving the problem after fault location. Table 1 shows a server system emergency drill log, including columns such as the drill date, drill personnel, and system version. The following information is filled in by sequence: Alarm step: alarm receipt time, alarm information, and screenshots; Location step: location completion time, location problem, and screenshots; Resolution step: emergency resolution completion time, emergency resolution description, and screenshots. The alarm time difference, location time difference, and resolution time difference are then calculated.

[0090] In some optional embodiments, the execution unit includes a fourth calculation module, a fifth calculation module, a sixth calculation module and a seventh calculation module, the fourth calculation module is used to obtain the total number of faults, and determine the number of alarms whose alarm time difference is less than the second preset threshold, calculate the ratio of the above-mentioned number of alarms to the above-mentioned total number of faults, and obtain the alarm score, calculate the product of the ratio of the above-mentioned number of alarms to the above-mentioned total number of faults and the above-mentioned alarm score, and obtain the above-mentioned alarm trigger score, wherein the above-mentioned alarm score represents the proportion of the fault alarm in the above-mentioned scheme score; the fifth calculation module is used to determine the number of positioning times whose positioning time difference is less than the third preset threshold, calculate the ratio of the above-mentioned number of positioning to the above-mentioned total number of faults, and obtain the positioning score, calculate the ratio of the above-mentioned number of positioning to the above-mentioned total number of faults, and obtain the positioning score, calculate the ratio of the above-mentioned number of positioning to the above-mentioned total number of faults, and obtain the positioning score. The product of the ratio of the above-mentioned total number of faults and the above-mentioned positioning score is used to obtain the above-mentioned positioning problem score, wherein the above-mentioned positioning score represents the proportion of fault positioning in the above-mentioned solution score; the sixth calculation module is used to obtain the number of located faults, determine the number of treatments whose treatment time difference is less than the fourth preset threshold, calculate the ratio of the above-mentioned treatment number to the above-mentioned number of located faults, and obtain the treatment score, calculate the product of the ratio of the above-mentioned treatment number to the above-mentioned number of located faults and the above-mentioned treatment score, and obtain the above-mentioned treatment problem score, wherein the above-mentioned treatment score represents the proportion of fault treatment in the above-mentioned solution score; the seventh calculation module is used to calculate the sum of the above-mentioned alarm trigger score, the above-mentioned positioning problem score and the above-mentioned treatment problem score to obtain the above-mentioned solution score. Through the above-mentioned steps, the device clarifies the efficiency requirements of the three stages of alarm response, fault location and fault treatment, and calculates the scores in combination with the scoring rules, making the evaluation of emergency response capabilities more objective and quantitative.

[0091] Specifically, the calculation of the alarm trigger score: First, obtain the total number of failures in the entire emergency drill process, which is a quantitative indicator that represents all possible failure conditions set in the drill. Then, determine the number of alarms whose alarm time difference is less than the second preset threshold, that is, those situations where the response speed is faster than the preset standard after the failure occurs. Calculate the ratio of the number of alarms to the total number of failures. This step evaluates the team's ability to respond quickly to failures. Next, multiply it by the alarm score. The alarm score here represents the proportion of fault alarms in the entire emergency drill scoring system. For example, the alarm score is 10 points. Finally, get the alarm trigger score, which reflects the efficiency and performance of the team in responding to alarms.

[0092] Calculating the Location Problem Score: In this step, we first obtain the total number of faults in the drill. Next, we find the number of locations where the location time difference is less than the third preset threshold, that is, the number of cases in which the team successfully located the fault within the specified time. By calculating the ratio of the number of locations to the total number of faults, we evaluate the team's efficiency in fault location. This is then multiplied by the location score, which represents the importance of fault location in the overall assessment. Finally, we obtain the location problem score, which measures the team's fault location capabilities and response speed. For example, a location score of 20 points is sufficient.

[0093] Calculating the Problem Handling Score: First, obtain the number of faults that need to be located during the drill, i.e., the number of located faults. Next, find the number of faults whose resolution time difference is less than the fourth preset threshold. This indicates that the team was able to complete fault handling within the preset time. By calculating the ratio of the number of resolved faults to the number of located faults, the team's efficiency and ability in fault handling are assessed. This is then multiplied by the resolution score, which reflects the weight of fault handling in the overall evaluation system. Finally, the Problem Handling Score is obtained, which reflects the team's performance during the fault handling phase. For example, a resolution score of 30 points is used. Determining the Final Solution Score: The three scores mentioned above—the alarm trigger score, the problem location score, and the problem handling score—are added together to obtain the solution score for the entire fault drill. This total score intuitively demonstrates the emergency response team's comprehensive performance in the three key aspects of rapid response, accurate location, and efficient resolution.

[0094] In some optional embodiments, the execution unit includes a determination module and an eighth calculation module, which can be implemented by the following steps: the determination module is used to determine the plan score and the total answer score, wherein the plan score represents the degree of compliance of the troubleshooting steps with the pre-set steps, and the total answer score is the score obtained by the user correctly answering the server fault problem; the eighth calculation module is used to calculate the sum of the solution score, the plan score and the total answer score to obtain the total score. The device comprehensively considers multiple dimensions of emergency drills through the calculation of the total score, not only evaluating the actual operational capabilities of emergency handling, but also considering the team's implementation of the plan and the members' mastery of theoretical knowledge, providing a comprehensive and in-depth assessment of emergency response capabilities.

[0095] Specifically, the plan score is determined by assessing the degree to which the troubleshooting steps align with the pre-defined emergency response plan. During the emergency drill, the team's troubleshooting setup is compared against the pre-defined plan. The plan score is designed to verify whether the team adhered to standard emergency procedures during the drill and the practicality of the plan. A high plan score indicates that the team was able to skillfully and accurately execute the plan, reflecting the plan's feasibility and the team's execution capabilities. The overall score is a test score of the user's knowledge points during the emergency drill, including identification of fault types, mastery of emergency response procedures, and understanding of troubleshooting strategies. This step tests team members' theoretical knowledge and practical skills to ensure the team's technical capabilities to handle various faults. A high overall score indicates that team members have a solid grasp of theoretical knowledge of emergency response, providing a solid theoretical foundation for actual troubleshooting. The overall score is calculated by adding the solution score, the plan score, and the overall score. The solution score reflects the team's efficiency and capabilities in actual troubleshooting drills, the contingency plan score verifies the effectiveness of the plan's execution, and the overall score reflects the team members' theoretical knowledge. By combining these three, the overall score comprehensively reflects the emergency response team's performance in all aspects of emergency handling, including rapid response, accurate execution of the plan, and in-depth theoretical knowledge. Table 2 shows a quantitative scoring table; the total score is used to calculate this overall score.

[0096] In some optional embodiments, the determination module includes an acquisition submodule and a first calculation submodule. The acquisition submodule is configured to acquire a preset handling plan, wherein the preset handling plan represents a plurality of historical fault resolution steps pre-set prior to the fault drill. The first calculation submodule is configured to calculate the ratio of the number of faults for which the fault resolution steps are identical to the historical fault resolution steps to the total number of faults, thereby obtaining the plan score. Through these steps, the device verifies the practicality and implementation effectiveness of the emergency plan, promotes the standardization and consistency of the team's emergency response capabilities, and thus provides an effective evaluation and improvement tool for improving overall emergency operation and maintenance capabilities.

[0097] Specifically, before the emergency drill begins, the system or team will pre-define a set of solutions (i.e., historical troubleshooting steps) based on historical fault data and expert experience. These solutions cover a variety of possible fault types and corresponding handling processes. These pre-determined solutions serve as guidelines for the team to follow when facing specific faults. They aim to improve the efficiency and success rate of fault handling through standardized and regularized emergency response processes. During the drill, the team will execute a series of troubleshooting steps based on the fault they encounter. The system records these actual troubleshooting steps and compares them with the historical troubleshooting steps in the pre-determined solutions. Specifically, the system counts the number of troubleshooting steps executed during the drill that exactly match the pre-determined solutions, then divides this number by the total number of faults encountered during the drill to calculate a ratio. This ratio reflects the degree to which the troubleshooting steps executed by the team during the drill align with the requirements of the emergency plan. Once this ratio is calculated, the system converts it into a plan score, which directly reflects the degree to which the team's emergency response process aligns with the pre-determined solutions. The higher the score, the more the emergency measures taken by the team during the drill are in line with historical experience and plans developed by experts, demonstrating the team's good grasp of the plan and its ability to execute it.

[0098] In some optional embodiments, the determination module further includes a second calculation submodule configured to conduct a fault drill in the aforementioned fault drill scenario and obtain multiple answer scores, calculate the average of these multiple answer scores, and obtain an overall answer score, where the answer score represents the score of each user's answer. Based on the user's answer scores, the device can identify areas where team members have knowledge gaps or misunderstandings, providing guidance for subsequent training and knowledge supplementation.

[0099] Specifically, the emergency drill platform simulates a series of scenarios related to possible server failures. These scenarios include hardware failures such as hard drive and memory failures, as well as software failures such as operating system anomalies, application crashes, and network outages. In these scenarios, users are asked to identify the failure type, analyze the cause, and propose solutions, or answer questions about troubleshooting procedures, standards, and best practices. During the drill, each user independently completes a series of theoretical knowledge tests, which may include multiple-choice questions, fill-in-the-blank questions, and short-answer questions. The user's answers are immediately scored by the system, based on factors such as the accuracy of the answer, the rationality of the solution, and familiarity with emergency response procedures. Each user receives a score based on their performance. The system collects all user scores and then averages them to obtain an overall score for the entire team. This average calculation is simple and intuitive, providing a comprehensive reflection of the team members' overall theoretical knowledge and mastery of troubleshooting knowledge.

[0100] In some optional embodiments, the execution unit includes an acquisition module, a ninth calculation module, and an execution module, wherein the acquisition module is used to acquire the total number of improvement measures and the total number of the improvement measures adopted, wherein the improvement measures are improvement measures for a plurality of the fault resolution steps; the ninth calculation module is used to calculate the ratio of the total number of the improvement measures adopted to the total number of the improvement measures to obtain an adoption score; and the execution module is used to update the corresponding fault resolution steps according to the adopted improvement measures when the adoption score is greater than a fifth preset threshold, and execute each of the fault resolution steps to resolve the fault of the server. By associating the adoption score with the improvement measures, the device promotes the regular optimization and updating of emergency plans and fault resolution steps, thereby improving the applicability and efficiency of the plans.

[0101] Specifically, the total number of improvement measures and the total number of adopted improvement measures are captured. After each emergency drill, the team reviews the drill, summarizes lessons learned, and proposes improvement measures. These improvement measures may include optimizing troubleshooting procedures, adjusting emergency response plans, and improving resource allocation. The system records the total number of proposed improvement measures and the number of those subsequently adopted and implemented. This step provides the basic data for calculating the adoption score. By calculating the ratio of the total number of adopted improvement measures to the total number of improvement measures, the system generates an adoption score. The adoption score reflects the team's attention to emergency drill feedback and the implementation rate of improvement measures. A high adoption score indicates that the team effectively learned from the drill and converted feedback into concrete improvement actions. When the adoption score exceeds the fifth preset threshold, it indicates a high adoption rate and that the team actively applies post-drill feedback to optimize troubleshooting procedures. Based on the adopted improvement measures, the system automatically or assists the team in updating troubleshooting procedures, ensuring that plans and emergency response strategies reflect the latest operational knowledge and best practices. Once the troubleshooting steps have been updated, the team will execute these optimized steps to resolve actual server failures, particularly those matching the target failure types in the drill. The updated steps are more efficient and precise, helping the team perform better in actual troubleshooting, reducing recovery time and improving system stability. Table 3 shows a statistical table for continuous improvement measures, used to record improvement measures.

[0102] The server failure processing device includes a processor and memory. The configuration unit, determination unit, and execution unit are stored as program units in the memory. The processor executes the program units stored in the memory to implement the corresponding functions. The modules are all located in the same processor; alternatively, the modules can be located in different processors in any combination.

[0103] The processor contains a kernel, which retrieves the corresponding program unit from the memory. One or more kernels can be configured, and the server system's emergency operation and maintenance capabilities can be quantified by adjusting kernel parameters.

[0104] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.

[0105] An embodiment of the present invention provides a computer-readable storage medium, which includes a stored program. When the program is executed, the device where the computer-readable storage medium is located is controlled to execute the server failure processing method.

[0106] Specifically, methods for handling server failures include:

[0107] Step S201: Obtain the fault type of the server fault, and configure a fault drill scenario according to the fault type, wherein the fault drill scenario at least includes the model of the server;

[0108] Step S202: Perform a fault drill in the fault drill scenario and obtain multiple fault resolution steps, and determine a time difference between each of the fault resolution steps, wherein the fault resolution step is a step for resolving the server fault, and the time difference represents the time consumed by each of the fault resolution steps.

[0109] Step S203: Determine a solution score for the fault drill based on each of the above time differences, and determine a total score based on the above solution score. At least when the above total score is greater than a first preset threshold and a target fault type fault occurs on the server, execute multiple fault resolution steps corresponding to the target fault type to resolve the server fault, wherein the target fault type is one of the above fault types.

[0110] An embodiment of the present invention provides an electronic device, including a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, at least the following steps are performed:

[0111] Step S201: Obtain the fault type of the server fault, and configure a fault drill scenario according to the fault type, wherein the fault drill scenario at least includes the model of the server;

[0112] Step S202: Perform a fault drill in the fault drill scenario and obtain multiple fault resolution steps, and determine a time difference between each of the fault resolution steps, wherein the fault resolution step is a step for resolving the server fault, and the time difference represents the time consumed by each of the fault resolution steps.

[0113] Step S203: Determine a solution score for the fault drill based on each of the above time differences, and determine a total score based on the above solution score. At least when the above total score is greater than a first preset threshold and a target fault type fault occurs on the server, execute multiple fault resolution steps corresponding to the target fault type to resolve the server fault, wherein the target fault type is one of the above fault types.

[0114] The devices in this article can be servers, PCs, PADs, mobile phones, etc.

[0115] The present application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the above method in each embodiment of the present application are implemented:

[0116] Step S201: Obtain the fault type of the server fault, and configure a fault drill scenario according to the fault type, wherein the fault drill scenario at least includes the model of the server;

[0117] Step S202: Perform a fault drill in the fault drill scenario and obtain multiple fault resolution steps, and determine a time difference between each of the fault resolution steps, wherein the fault resolution step is a step for resolving the server fault, and the time difference represents the time consumed by each of the fault resolution steps.

[0118] Step S203: Determine a solution score for the fault drill based on each of the above time differences, and determine a total score based on the above solution score. At least when the above total score is greater than a first preset threshold and a target fault type fault occurs on the server, execute multiple fault resolution steps corresponding to the target fault type to resolve the server fault, wherein the target fault type is one of the above fault types.

[0119] Obviously, those skilled in the art will appreciate that the various modules or steps of the present invention described above can be implemented using a general-purpose computing device, can be centralized on a single computing device, or can be distributed across a network of multiple computing devices. They can be implemented using program code executable by the computing device, and thus, can be stored in a storage device and executed by the computing device. In some cases, the steps shown or described herein can be performed in a different order than that shown, or can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.

[0120] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0121] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0122] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0123] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0124] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0125] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0126] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0127] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0128] From the above description, it can be seen that the above embodiments of the present application achieve the following technical effects:

[0129] 1) In the server fault handling method of the present application, a fault type is obtained, and a fault drill scenario is configured based on the fault type. The fault type represents the type of server fault, and the fault drill scenario includes at least the server model. A fault drill is performed in the fault drill scenario and multiple fault resolution steps are obtained. The time difference of each fault resolution step is determined. The fault resolution step is a step for resolving the server fault, and the time difference represents the time consumed by each fault resolution step. A solution score for the fault drill is determined based on each time difference, and a total score is determined based on the solution score. At least when the total score is greater than a first preset threshold and the server has a fault of the fault type, the multiple fault resolution steps are executed to resolve the server fault. Compared with the prior art, in which the emergency operation and maintenance capabilities of the server system cannot be quantified, resulting in the inability to discover and address problems in the financial system, thereby causing economic losses, the present application scores the fault resolution steps through the above scheme to quantify the problem-solving capabilities. The quantified total score is then used to determine whether to continue using the fault resolution steps to handle the fault. If the total score is greater than the first preset threshold, the fault resolution steps are continued to handle the fault, thereby promptly discovering and addressing problems in the server system. Therefore, it can solve the problem that the emergency operation and maintenance capabilities of the server system cannot be quantified, resulting in the inability to discover and deal with problems in the financial system, thereby causing economic losses, and achieve the effect of reducing economic losses.

[0130] 2) In the server fault handling device of the present application, a fault type is obtained, and a fault drill scenario is configured based on the fault type. The fault type indicates the type of server fault, and the fault drill scenario includes at least the server model. A fault drill is performed in the fault drill scenario and multiple fault resolution steps are obtained. The time difference of each fault resolution step is determined. The fault resolution step is a step for resolving the server fault, and the time difference indicates the time consumed by each fault resolution step. A solution score for the fault drill is determined based on each time difference, and a total score is determined based on the solution score. At least when the total score is greater than a first preset threshold and the server has a fault of the fault type, the multiple fault resolution steps are executed to resolve the server fault. Compared with the prior art, in which the emergency operation and maintenance capabilities of the server system cannot be quantified, resulting in the inability to discover and address problems in the financial system, thereby causing economic losses, the present application scores the fault resolution steps through the above scheme to quantify the problem-solving capabilities. The quantified total score is then used to determine whether to continue using the fault resolution steps to handle the fault. If the total score is greater than the first preset threshold, the fault resolution steps are continued to handle the fault, thereby promptly discovering and addressing problems in the server system. Therefore, it can solve the problem that the emergency operation and maintenance capabilities of the server system cannot be quantified, resulting in the inability to discover and deal with problems in the financial system, thereby causing economic losses, and achieve the effect of reducing economic losses.

[0131] The above description is merely a preferred embodiment of the present application and is not intended to limit the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.

Claims

1. A method for handling server failure, characterized in that: include: Obtaining a fault type of a fault occurring on the server, and configuring a fault drill scenario according to the fault type, wherein the fault drill scenario at least includes the model of the server; Performing a fault drill in the fault drill scenario and obtaining multiple fault resolution steps, and determining a time difference between each of the fault resolution steps, wherein the fault resolution step is a step of resolving the server fault, and the time difference represents the time consumed by each of the fault resolution steps; A solution score for the fault drill is determined based on each of the time differences, and a total score is determined based on the solution score. At least when the total score is greater than a first preset threshold and a target fault type fault occurs on the server, multiple fault resolution steps corresponding to the target fault type are executed to resolve the server fault, wherein the target fault type is one of the fault types.

2. The method for handling server failure according to claim 1, characterized in that: Determine the time difference for each of the described troubleshooting steps, including: Obtaining an alarm receiving time and an alarm issuing time, and calculating a difference between the alarm receiving time and the alarm issuing time to obtain an alarm time difference, wherein the alarm issuing time indicates the time when the fault alarm is issued, and the alarm receiving time indicates the time when the fault alarm is received; Obtaining a fault location time, and calculating a difference between the fault location time and the alarm reception time to obtain a location time difference, wherein the fault location time represents the time when the fault location is received; Obtain the fault handling time, and calculate the difference between the fault handling time and the fault location time to obtain the handling time difference, wherein the fault handling time represents the time when the fault handling completion is received.

3. The method for handling server failure according to claim 2, characterized in that: Determining a solution score for the fault drill according to each of the time differences includes: Obtaining the total number of faults, determining the number of alarms for which the alarm time difference is less than a second preset threshold, calculating a ratio of the number of alarms to the total number of faults, obtaining an alarm score, and calculating the product of the ratio of the number of alarms to the total number of faults and the alarm score to obtain an alarm trigger score, wherein the alarm score represents a proportion of the fault alarm in the solution score; Determining the number of positioning operations for which the positioning time difference is less than a third preset threshold, calculating a ratio of the number of positioning operations to the total number of faults, obtaining a positioning score, and calculating the product of the ratio of the number of positioning operations to the total number of faults and the positioning score to obtain a positioning problem score, wherein the positioning score represents the proportion of fault positioning in the solution score; Obtaining the number of located faults, determining the number of resolved faults for which the resolution time difference is less than a fourth preset threshold, calculating a ratio of the number of resolved faults to the number of located faults, obtaining a resolution score, and calculating the product of the ratio of the number of resolved faults to the number of located faults and the resolution score to obtain a resolution score, wherein the resolution score represents a proportion of the fault resolution in the solution score; The sum of the alarm trigger score, the positioning problem score, and the handling problem score is calculated to obtain the solution score.

4. The method for handling server failure according to claim 1, wherein: The overall score is determined based on the solution score, including: Determine a plan score and a total answer score, wherein the plan score represents the degree to which the troubleshooting steps conform to the pre-set steps, and the total answer score is the score obtained by the user correctly answering the server troubleshooting question; The sum of the solution score, the plan score and the total answer score is calculated to obtain the total score.

5. The method for handling server failure according to claim 4, characterized in that: Determine the plan score, including: Obtaining a preset processing solution, wherein the preset processing solution represents a plurality of historical fault resolution steps preset before the fault drill; The ratio of the number of faults for which the fault resolution steps are the same as the historical fault resolution steps to the total number of faults is calculated to obtain the plan score.

6. The method for handling server failure according to claim 4, characterized in that: Determine the overall score for your questions, including: A fault drill is performed in the fault drill scenario and multiple answer scores are obtained, and an average of the multiple answer scores is calculated to obtain a total answer score, wherein the answer score is a score for each user's answer.

7. The method for handling server failure according to claim 1, wherein: At least when the total score is greater than a first preset threshold and a fault of a target fault type occurs on the server, executing a plurality of fault resolution steps corresponding to the target fault type to resolve the server fault includes: Obtaining a total number of improvement measures and a total number of adopted improvement measures, wherein the improvement measures are measures for improving a plurality of the fault resolution steps; Calculating the ratio of the total number of the adopted improvement measures to the total number of the improvement measures to obtain an adoption score; In a case where the adoption score is greater than a fifth preset threshold, the corresponding fault resolution steps are updated according to the adopted improvement measures, and each fault resolution step is executed to resolve the server fault.

8. A server failure processing device, characterized in that: include: a configuration unit, configured to obtain a fault type of a fault occurring on the server, and configure a fault drill scenario according to the fault type, wherein the fault drill scenario at least includes a model of the server; a determining unit, configured to perform a fault drill in the fault drill scenario and obtain a plurality of fault resolution steps, and determine a time difference between each of the fault resolution steps, wherein the fault resolution step is a step of resolving the fault of the server, and the time difference represents a time consumed by each of the fault resolution steps; an execution unit, configured to determine a solution score for the fault drill based on each of the time differences, and to determine a total score based on the solution score, and to execute a plurality of fault resolution steps corresponding to the target fault type to resolve the server fault, at least when the total score is greater than a first preset threshold and a fault of a target fault type occurs on the server, wherein the target fault type is one of the fault types.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the method for handling server failure according to any one of claims 1 to 7.

10. An electronic device, characterized in that: include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include a method for processing a server failure according to any one of claims 1 to 7.