Fault handling method and system based on emergency plan
By generating fault listeners and handling scripts in the cloud computing environment and automating the parsing of emergency plan templates, the problem of poor flexibility of traditional methods in the cloud computing environment is solved, and timely and efficient handling of various faults is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CCB FINTECH CO LTD
- Filing Date
- 2021-11-22
- Publication Date
- 2026-05-01
AI Technical Summary
Traditional fault recovery methods are inflexible in cloud computing environments and cannot adapt to complex fault scenarios, resulting in insufficient flexibility and timeliness in the fault handling methods of operation and maintenance tools and cloud computing platforms.
Based on the emergency plan, multiple fault listeners are generated. After a fault occurs, a fault handling script is generated. The emergency plan template is parsed using automated text parsing, and the fault handling script is executed to restore the fault, including fault recovery operations and multiple verifications, and a fault handling log is generated.
It enables timely handling of various faults, has scalability and versatility, a high degree of automation, and a short fault handling cycle, thus improving the flexibility and accuracy of fault handling.
Smart Images

Figure CN114064341B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of operation and maintenance technology, and in particular to a fault handling method, system, equipment, medium and product based on an emergency plan. Background Technology
[0002] In cloud environments, cloud-based applications often experience a series of failures, such as process failures. Some commonly used operation and maintenance tools and cloud computing platforms provide simple failure recovery functions based on processes.
[0003] However, traditional fault recovery methods are process-based and have relatively simple targets. In real-world cloud computing environments, faults often come in various types, and the corresponding fault handling methods are more complex. Fault handling methods for processes and other similar objects are often inflexible and not suitable for scenarios with changing circumstances. Summary of the Invention
[0004] In view of the above issues, this disclosure provides fault handling methods, systems, equipment, media and program products based on emergency plans.
[0005] According to a first aspect of this disclosure, a fault handling method based on an emergency plan is provided, comprising: generating multiple fault listeners according to a preset emergency plan, the fault listeners being used to detect fault scenarios; when the fault listeners detect a fault occurring, generating a fault handling script according to the preset emergency plan; and executing the fault handling script to restore the fault.
[0006] According to embodiments of this disclosure, generating a fault handling script based on the preset emergency plan includes:
[0007] The preset emergency plan is scanned to generate a fault handling script.
[0008] According to embodiments of this disclosure, the step of scanning the preset emergency plan to generate a fault handling script includes:
[0009] The preset emergency plan is parsed to determine the trigger actions and scenario verification after fault recovery;
[0010] Generate emergency recovery statement code based on the triggered execution action;
[0011] Generate fault recovery verification statement code based on the scenario verification after fault recovery;
[0012] A fault handling script is generated based on the recovery statement code and the fault recovery verification statement code.
[0013] According to embodiments of this disclosure, executing the fault handling script to recover from the fault includes:
[0014] Execute emergency recovery statement code to complete the fault recovery operation;
[0015] Execute the fault recovery verification statement code a preset number of times to verify the fault scenario multiple times.
[0016] According to embodiments of this disclosure, executing the fault recovery verification statement code a preset number of times to perform multiple checks on the fault scenario includes:
[0017] Execute the fault recovery verification statement code a preset number of times and obtain the verification results;
[0018] If the test result is determined to be a test failure and the number of test failures is equal to the preset number, then the fault handling is determined to be a failure.
[0019] Re-perform fault scenario detection and / or issue alarm messages.
[0020] According to embodiments of this disclosure, generating multiple fault listeners based on a preset emergency plan includes:
[0021] The activation scenario and triggering conditions of the emergency plan are determined according to the preset emergency plan;
[0022] Multiple detection scripts are generated based on the activation scenario and triggering conditions of the emergency plan, wherein the detection scripts correspond to the activation scenario of the emergency plan.
[0023] According to embodiments of this disclosure, after executing the fault handling script to recover from the fault, the method further includes:
[0024] Generate a fault handling log, which includes the procedure steps before and during the fault, as well as the fault recovery status.
[0025] The second aspect of this disclosure provides a fault handling system based on an emergency plan, including: a fault monitoring module, used to generate multiple fault monitors according to a preset emergency plan, wherein the fault monitors are used to detect fault scenarios;
[0026] The fault handling procedure generation module is used to generate a fault handling script based on the preset emergency plan; and
[0027] The fault handling procedure execution module is used to execute the fault handling script to restore the fault.
[0028] A third aspect of this disclosure provides an electronic device comprising: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors perform the aforementioned fault handling method based on an emergency plan.
[0029] A fourth aspect of this disclosure also provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, cause the processor to perform the aforementioned fault handling method based on the contingency plan.
[0030] The fifth aspect of this disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described fault handling method based on an emergency plan.
[0031] The fault handling method based on emergency plans provided in this disclosure generates a monitor according to a preset emergency plan to acquire faults generated by the system in real time. When the fault monitor detects a fault, it generates a fault handling script according to the preset emergency plan and executes the fault handling script to restore the fault. The fault handling script is matched with the fault type to realize timely handling of various faults. The emergency plan template with a specified format is parsed by automated text parsing, so it has a certain degree of scalability and universality, is relatively dynamic and flexible, has a high degree of automation and accuracy, and a short fault handling cycle. Attached Figure Description
[0032] The foregoing contents, as well as other objects, features, and advantages of this disclosure, will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:
[0033] Figure 1 A flowchart illustrating a fault handling process based on an emergency plan according to an embodiment of the present disclosure is shown schematically.
[0034] Figure 2 This schematically illustrates a system architecture diagram that can be used in a fault handling method based on an emergency response plan according to embodiments of the present disclosure;
[0035] Figure 3 A flowchart illustrating a fault handling method based on an emergency plan according to an embodiment of the present disclosure is shown schematically.
[0036] Figure 4 The flowchart schematically illustrates another fault handling method based on an emergency plan according to an embodiment of the present disclosure;
[0037] Figure 5 A schematic diagram illustrates a structural block diagram of a fault handling system based on an emergency plan according to an embodiment of the present disclosure; and
[0038] Figure 6 A block diagram schematically illustrates an electronic device suitable for implementing a fault handling method based on an emergency plan, according to an embodiment of the present disclosure. Detailed Implementation
[0039] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.
[0040] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0041] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0042] When using expressions such as "at least one of A, B, and C", they should generally be interpreted in accordance with the meaning that is commonly understood by a person skilled in the art (e.g., "a system having at least one of A, B, and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B, and C, etc.).
[0043] With the continuous development of the Internet and cloud computing, cloud applications are becoming more and more numerous, and the requirements for cloud service quality are also getting higher and higher. This puts forward higher requirements for operation and maintenance. In one example, when a cloud application fails, common operation and maintenance tools and cloud computing platforms will provide some simple fault recovery functions. These process-based fault recovery functions have relatively simple targets. When faced with complex faults, manual intervention by operation and maintenance personnel is required, and it is impossible to deal with the fault in a timely manner.
[0044] To address the aforementioned technical issues, embodiments of this disclosure provide a fault handling method based on an emergency plan. The method includes: generating multiple fault listeners according to a preset emergency plan, wherein the fault listeners are used to detect fault scenarios; when a fault listener detects a fault occurring, generating a fault handling script according to the preset emergency plan; and executing the fault handling script to restore the fault.
[0045] Figure 1 A flowchart illustrating a fault handling process based on an emergency plan according to an embodiment of the present disclosure is shown. Figure 2This diagram schematically illustrates a system architecture that can be used in a fault handling method based on an emergency response plan according to embodiments of this disclosure. It should be noted that... Figure 1 The flowchart shown and Figure 2 The system architecture shown is merely an example of application scenarios and system architectures that can be used in the embodiments of this disclosure, to help those skilled in the art understand the technical content of this disclosure, but does not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments, or scenarios. It should be noted that the fault handling method and system based on emergency plans provided in the embodiments of this disclosure can be used in the field of cloud computing operation and maintenance technology, related aspects of the financial field, and can also be used in any field other than the financial field. The application fields of the fault handling method and system based on emergency plans provided in the embodiments of this disclosure are not limited.
[0046] like Figure 2 As shown, the system provided in this embodiment is divided into an emergency plan scanning module, a fault handling module, and a fault log analysis module, as follows: Figure 1 As shown, the emergency plan scanning module is used to scan and parse emergency plans, generating fault detection and fault handling procedures. When a fault is detected, the emergency plan document is scanned according to the fault type to determine the fault handling procedure that matches the fault. The fault handling module loads and executes the fault handling procedure. After the fault is handled, the fault handling completion verification procedure is run. This verification procedure is also generated by parsing the emergency plan through the emergency plan scanning module. The verification result determines whether the fault has been handled successfully. If the verification is successful, it indicates that the current fault has been handled successfully. If the verification fails multiple times in a row, it is determined that the current fault handling has failed, and the fault will be re-detected and an alarm message will be issued. The entire process of fault handling is recorded by the fault handling logger, which facilitates the analysis of faults by maintenance personnel.
[0047] The following will be based on Figure 1 and Figure 2 The described fault handling process, through Figures 3-4 The fault handling method based on the emergency plan in the disclosed embodiments is described in detail.
[0048] Figure 3 A flowchart illustrating a fault handling method based on an emergency plan according to an embodiment of the present disclosure is shown schematically.
[0049] like Figure 3 As shown, the fault handling method based on the emergency plan in this embodiment includes operations S310 to S330, and the fault handling method can be executed by a server or other computing device.
[0050] When operating S310, multiple fault listeners are generated according to the preset emergency plan.
[0051] In one example, the preset emergency plan is structured according to rules, meaning the text of the emergency plan is formatted according to defined constraints or tables. The preset emergency plan in this embodiment is divided into three parts:
[0052] a) Scenario when the emergency plan occurs, determining specific inspection commands or monitoring conditions based on the emergency scenario; b) Operations to be performed when an emergency occurs, detailing specific operation commands or program instructions; c) Verification statements after performing emergency operations, mainly used to verify the corresponding fault handling results, detailing specific verification commands and expected assertions.
[0053] The structural rules of the emergency response plan can be stored in an Excel spreadsheet or database table in the form of a specification. The program then reads and parses the data in the corresponding Excel spreadsheet and database table to perform command parsing. In this operation, the first part of the preset emergency response plan is parsed, and the statement code at the time of the emergency response is captured to generate multiple listeners. Fault listeners are used to detect fault scenarios. Specifically, the observer design pattern in programming is used to listen for faults. Once a fault scenario is triggered, the corresponding listener will execute the fault recovery operation. The number of listeners corresponds to the emergency scenario (fault scenario), meaning different emergency scenarios correspond to different listeners, achieving coverage of multiple fault scenarios in the cloud computing environment, real-time fault detection, and ensuring the stability of system operation.
[0054] When operating the S320, once the fault listener detects a fault, it generates a fault handling script based on the preset emergency plan.
[0055] In one example, when the fault listener detects a fault, it will perform a fault recovery operation. Specifically, a fault handling script is generated according to the preset emergency plan. The fault handling script will be different depending on the type of fault. Compared with the process-based fault restart operation in the prior art, the fault handling script provided in this embodiment is generated based on the emergency scenario in the emergency plan, that is, based on the fault type. Therefore, the fault handling method provided in this embodiment can handle faults in a variety of complex scenarios.
[0056] In operation S330, the fault handling script is executed to restore the fault.
[0057] In one example, the fault handling script generated by operation S320 is executed to recover from the fault. The fault handling script also includes a verification procedure for the fault handling result, which is used to verify the fault handling result after the fault handling is completed to ensure that the fault is successfully handled. When the verification result indicates that the fault handling has failed, the fault will be re-detected and an alarm message will be issued.
[0058] The fault handling method based on emergency plans provided in this disclosure generates a monitor according to a preset emergency plan to acquire faults generated by the system in real time. When the fault monitor detects a fault, it generates a fault handling script according to the preset emergency plan and executes the fault handling script to restore the fault. The fault handling script is matched with the fault type to realize timely handling of various faults. The emergency plan template in the prescribed format is parsed by automated text parsing, so it has a certain degree of scalability and universality, is relatively dynamic and flexible, has a high degree of automation and accuracy, and a short fault handling cycle.
[0059] Figure 4 A flowchart illustrating another fault handling method based on an emergency plan according to an embodiment of this disclosure is shown schematically. Figure 4 As shown, this includes operations S410 to S450.
[0060] When operating the S410, multiple fault listeners are generated according to the preset emergency plan.
[0061] According to an embodiment of this disclosure, the activation scenario and triggering conditions of an emergency plan are determined based on a preset emergency plan; multiple detection scripts are generated based on the activation scenario and triggering conditions of the emergency plan, wherein the detection scripts correspond to the activation scenario of the emergency plan.
[0062] In one example, the emergency response plan is formed by operations and maintenance personnel according to common fault types and preset structural rules. An example of an emergency response plan template is shown in Table 1 below:
[0063] Table 1 Emergency Response Plan Template
[0064]
[0065] Because the emergency response plan is updated in a timely manner and has a high real-time processing capability, the fault handling procedure can be expanded dynamically as the emergency response plan items expand without requiring program modification. The program can adaptively capture items in the emergency response plan and automatically parse them as the number of emergency response plan items increases, thus shortening the system program's feedback cycle.
[0066] When operating the S420, scan the preset emergency plan to generate a fault handling script.
[0067] According to an embodiment of this disclosure, operation S420 specifically includes the following operations:
[0068] In operation S421, the preset emergency plan is parsed to determine the trigger action and the scenario verification after fault recovery; in operation S422, emergency recovery statement code is generated based on the trigger action; in operation S423, fault recovery verification statement code is generated based on the scenario verification after fault recovery; in operation S424, fault handling script is generated based on the recovery statement code and the fault recovery verification statement code.
[0069] In one example, by parsing the text of the second and third parts of the preset emergency plan, the trigger action corresponding to the current fault and the scenario verification after the fault is recovered can be determined. Emergency recovery statement code and fault recovery verification statement code are generated according to the trigger action and the scenario verification after the fault is recovered, respectively. The two sets of codes are then used to generate a fault handling script in a preset order.
[0070] When operating S430, execute emergency recovery statement code to complete the fault recovery operation.
[0071] When operating S440, execute the fault recovery verification statement code a preset number of times to verify the fault scenario multiple times.
[0072] According to the embodiments of this disclosure, the fault recovery verification statement code is executed a preset number of times and the verification result is obtained; if the verification result is determined to be a verification failure and the number of verification failures is equal to the preset number, the fault handling is determined to be a failure; the fault scenario detection is re-performed and / or an alarm message is issued.
[0073] In one example, the emergency recovery statement code in the fault handling script is executed to complete the fault recovery operation. Afterwards, to ensure successful fault handling, the recovery verification statement code in the fault handling script is also executed to verify the fault scenario. Considering the system's instability, multiple verifications are required to ensure the accuracy of the verification results. Specifically, the fault recovery verification statement code is executed multiple times, resulting in multiple verification results. If multiple verification results are all failures, it indicates that the current fault handling has failed and the fault has not been resolved. The fault needs to be handled again. This could involve re-detecting the fault scenario to determine if the fault type is correct, or sending an alarm message to the operations and maintenance center to await manual intervention. If any of the multiple verification results are successful, it indicates that the current fault handling has been successful.
[0074] When operating the S450, a fault handling log is generated. The fault handling log includes the procedure steps before and during the fault, as well as the fault recovery status.
[0075] In one example, the system records the handling operations at each stage of the fault to generate a fault handling log, which includes the program steps before the fault occurs, when the fault occurs, and the recovery status after the fault occurs (including operations S410 to S440). This allows maintenance personnel to analyze the fault later, and the log has high traceability.
[0076] Based on the above-mentioned emergency response plan-based fault handling method, this disclosure also provides an emergency response plan-based fault handling system. The following will combine... Figure 5 The device is described in detail.
[0077] Figure 5 A schematic diagram of a fault handling system based on an emergency plan according to an embodiment of the present disclosure is shown.
[0078] like Figure 5 As shown, the fault handling system 500 based on the emergency plan in this embodiment includes a fault monitoring module 510, a fault handling program generation module 520, and a fault handling program execution module 530.
[0079] The fault monitoring module 510 is used to generate multiple fault listeners according to a preset emergency plan. The fault listeners are used to detect fault scenarios. In one embodiment, the fault monitoring module 510 can be used to perform the operation S210 described above, which will not be repeated here.
[0080] The fault handling procedure generation module 520 is used to generate a fault handling script based on a preset emergency plan. In one embodiment, the fault handling procedure generation module 520 can be used to execute the operation S220 described above, which will not be repeated here.
[0081] The fault handling procedure execution module 530 is used to execute the fault handling script to restore the fault. In one embodiment, the fault handling procedure execution module 530 can be used to execute the operation S230 described above, which will not be repeated here.
[0082] According to embodiments of this disclosure, any plurality of modules in the fault monitoring module 510, the fault handling procedure generation module 520, and the fault handling procedure execution module 530 may be combined into one module, or any one of these modules may be split into multiple modules. Alternatively, at least a portion of the functionality of one or more of these modules may be combined with at least a portion of the functionality of other modules and implemented in one module. According to embodiments of this disclosure, at least one of the fault monitoring module 510, the fault handling procedure generation module 520, and the fault handling procedure execution module 530 may be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging circuitry, or implemented in any one of software, hardware, and firmware methods, or in a suitable combination of any of these methods. Alternatively, at least one of the fault monitoring module 510, the fault handling program generation module 520, and the fault handling program execution module 530 may be implemented at least partially as a computer program module, which can perform corresponding functions when the computer program module is run.
[0083] Figure 6 A block diagram schematically illustrates an electronic device suitable for implementing a fault handling method based on an emergency plan, according to an embodiment of the present disclosure.
[0084] like Figure 6 As shown, an electronic device 900 according to an embodiment of the present disclosure includes a processor 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage portion 908 into a random access memory (RAM) 903. The processor 901 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 901 may also include onboard memory for caching purposes. The processor 901 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0085] RAM 903 stores various programs and data required for the operation of electronic device 900. Processor 901, ROM 902, and RAM 903 are interconnected via bus 904. Processor 901 performs various operations of the method flow according to embodiments of the present disclosure by executing programs in ROM 902 and / or RAM 903. It should be noted that programs may also be stored in one or more memories other than ROM 902 and RAM 903. Processor 901 may also perform various operations of the method flow according to embodiments of the present disclosure by executing programs stored in one or more memories.
[0086] According to embodiments of this disclosure, the electronic device 900 may further include an input / output (I / O) interface 905, which is also connected to a bus 904. The electronic device 900 may also include one or more of the following components connected to the I / O interface 905: an input section 906 including a keyboard, mouse, etc.; an output section 907 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN card, modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the I / O interface 905 as needed. A removable medium 911, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 910 as needed so that computer programs read from it can be installed into the storage section 908 as needed.
[0087] This disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs that, when executed, implement the method according to the embodiments of this disclosure.
[0088] According to embodiments of this disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, such as including, but not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of this disclosure, the computer-readable storage medium may include ROM 902 and / or RAM 903 and / or one or more memories other than ROM 902 and RAM 903 described above.
[0089] Embodiments of this disclosure also include a computer program product comprising a computer program containing program code for performing the methods shown in the flowchart. When the computer program product is run on a computer system, the program code enables the computer system to implement the fault handling method based on the emergency plan provided in the embodiments of this disclosure.
[0090] When the computer program is executed by the processor 901, it performs the functions defined in the system / apparatus of this disclosure embodiments. According to embodiments of this disclosure, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0091] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and downloaded and installed via the communication section 909, and / or installed from a removable medium 911. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.
[0092] In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 909, and / or installed from the removable medium 911. When the computer program is executed by the processor 901, it performs the functions defined in the system of this disclosure embodiment. According to embodiments of this disclosure, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0093] According to embodiments of this disclosure, program code for executing the computer programs provided in embodiments of this disclosure can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C", or similar programming languages. The program code can execute entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0094] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0095] Those skilled in the art will understand that the features described in the various embodiments and / or claims of this disclosure can be combined or combined in various ways, even if such combinations or combinations are not explicitly described in this disclosure. In particular, the features described in the various embodiments and / or claims of this disclosure can be combined or combined in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.
[0096] The embodiments of this disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of this disclosure. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. The scope of this disclosure is defined by the appended claims and their equivalents. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of this disclosure, and all such substitutions and modifications should fall within the scope of this disclosure.
Claims
1. A fault handling method based on an emergency plan, the method comprising: Multiple fault listeners are generated according to the preset emergency plan. The fault listeners are used to detect fault scenarios. The preset emergency plan is obtained through rule structuring. The preset emergency plan is a structured text that includes three parts: the start scenario and triggering conditions, the triggering action, and the verification statement after the emergency operation is performed. The structural rules of the preset emergency plan are stored in the database table in the form of a specification. When the fault listener detects a fault, it generates a fault handling script based on the preset emergency plan; and Execute the fault handling script to recover from the fault. The generation of multiple fault listeners according to the preset emergency plan includes: The activation scenario and triggering conditions of the emergency plan are determined according to the preset emergency plan; Multiple detection scripts are generated based on the activation scenario and triggering conditions of the emergency plan, wherein the detection scripts correspond to the activation scenario of the emergency plan; The step of generating a fault handling script based on the preset emergency plan includes: The preset emergency plan is parsed to determine the trigger actions and scenario verification after fault recovery; Generate emergency recovery statement code based on the triggered execution action; Generate fault recovery verification statement code based on the scenario verification after fault recovery; A fault handling script is generated based on the recovery statement code and the fault recovery verification statement code.
2. The method according to claim 1, characterized in that, The process of executing the fault handling script to recover from the fault includes: Execute emergency recovery statement code to complete the fault recovery operation; Execute the fault recovery verification statement code a preset number of times to verify the fault scenario multiple times.
3. The method according to claim 2, characterized in that, The fault recovery verification statement code, which executes a preset number of times to perform multiple checks on the fault scenario, includes: Execute the fault recovery verification statement code a preset number of times and obtain the verification results; If the test result is determined to be a test failure and the number of test failures is equal to the preset number, then the fault handling is determined to be a failure. Re-perform fault scenario detection and / or issue alarm messages.
4. The method according to any one of claims 1 to 3, characterized in that, After executing the fault handling script to recover from the fault, the method further includes: Generate a fault handling log, which includes the procedure steps before and during the fault, as well as the fault recovery status.
5. A fault handling system based on an emergency plan, comprising: The fault monitoring module is used to generate multiple fault listeners according to the preset emergency plan. The fault listeners are used to detect fault scenarios. The preset emergency plan is obtained through rule structuring. The preset emergency plan is a structured text that includes three parts: the start scenario and triggering conditions, the triggering action, and the verification statement after the emergency operation is performed. The structural rules of the emergency plan are stored in the database table in the form of a specification. The fault handling procedure generation module is used to generate a fault handling script based on the preset emergency plan; and The fault handling procedure execution module is used to execute the fault handling script to restore the fault. The fault monitoring module is also used to determine the activation scenario and triggering conditions of the emergency plan according to the preset emergency plan; and to generate multiple detection scripts according to the activation scenario and triggering conditions of the emergency plan, wherein the detection scripts correspond to the activation scenario of the emergency plan. The fault handling procedure generation module is also used to scan the preset emergency plan to generate a fault handling script; to parse the preset emergency plan to determine the triggering action and the scenario verification after fault recovery; to generate emergency recovery statement code based on the triggering action; to generate fault recovery verification statement code based on the scenario verification after fault recovery; and to generate a fault handling script based on the recovery statement code and the fault recovery verification statement code.
6. An electronic device, comprising: One or more processors; Storage device for storing one or more programs. Wherein, when the one or more programs are executed by the one or more processors, the one or more processors perform the method according to any one of claims 1-4.
7. A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, cause the processor to perform the method according to any one of claims 1-4.
8. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-4.
Citation Information
Patent Citations
Fault recovery method, computer equipment and storage medium
CN111897671A
Memory, accident emergency plan generation method, device and equipment
CN113094425A
Method for automatically testing, positioning and repairing a fault and storage medium
CN113656323A