Fault drill method, device, equipment and computer storage medium

By introducing automated fault drills and recovery processes into the financial system, using drill platforms, virtual machines, chaos engineering tools, intelligent monitoring systems and standard operating processes, the problem of low efficiency of fault drills in the existing technology has been solved, and efficient and automated fault drills and recovery has been achieved.

CN110308969BActive Publication Date: 2025-05-30WEBANK (CHINA)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201910570965.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-06-26
Publication Date
2025-05-30
Estimated Expiration
2039-06-26

AI Technical Summary

Technical Problem

The existing banking system fault drills are inefficient and rely on manual failure simulation and recovery, which is not efficient.

Method used

A fault drill method is proposed, using drill platform, virtual machines, chaos engineering tools, intelligent monitoring systems and standard operating processes to automate fault drills and recovery processes to realize closed-loop drills.

Benefits of technology

By automating the fault drill and recovery process, the risk of manual participation is reduced, the duration of failure impact is shortened, and the efficiency of financial system fault drills is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN110308969B_ABST
    Figure CN110308969B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of financial technology (Fintech), and discloses a fault drill method, which includes: when receiving a fault drill scenario instruction, controlling the drill platform to obtain an instance of a target VM in each VM based on the fault drill scenario instruction, and sending the instance and the fault drill scenario instruction to a chaos engineering tool through the drill platform; performing a drill on the instance through the chaos engineering tool, and monitoring whether the instance has completed the drill through IMS; if the drill is completed, sending an alarm message to the drill platform through IMS, and controlling the drill platform to obtain a recovery process in the SOP; sending the recovery process to the chaos engineering tool through the drill platform, performing a recovery on the instance through the chaos engineering tool, and sending a fault recovery message to the drill platform after the recovery is completed. The present invention also discloses a fault drill device, equipment and a computer storage medium. The present invention improves the efficiency of fault drills in the financial system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of financial technology (Fintech), and particularly to a method, device, equipment and computer storage medium for fault drill Background Art

[0002] With the development of computer technology, more and more technologies (big data, distributed, blockchain, artificial intelligence, etc.) are applied in the financial field. The traditional financial industry is gradually transforming into financial technology (Fintech). However, due to the security and real-time requirements of the financial industry, higher requirements are also put forward for technologies. For example, in order to maximize the resilience of the financial system (such as a banking system) to various emergencies, users will conduct fault drills on the banking system. However, most of the existing fault drills for banking systems are that the operation and maintenance personnel create fault scenarios manually according to business scenarios or host resources. After the operation and maintenance development team cooperates to troubleshoot problems in fixed scenarios, a recovery plan is formulated. The entire drill process from fault simulation to fault recovery depends on manual operation, and the efficiency is very low. Therefore, how to improve the efficiency of fault drills for financial systems has become an urgent technical problem to be solved at present. Summary of the Invention

[0003] The main purpose of the present invention is to provide a method, device, equipment and computer storage medium for fault drill, aiming to improve the efficiency of fault drills for financial systems.

[0004] To achieve the above object, the present invention provides a method for fault drill. The method for fault drill is applied to a fault drill system. The fault drill system includes: a drill platform, a plurality of virtual machines, a chaos engineering tool, an intelligent monitoring system and a standard operation process. The method for fault drill includes the following steps:

[0005] When receiving a fault drill scenario instruction, controlling the drill platform to obtain an instance of a target virtual machine in each of the virtual machines based on the fault drill scenario instruction, and sending the instance and the fault drill scenario instruction to the chaos engineering tool through the drill platform;

[0006] Based on the fault drill scenario instruction, conducting a drill on the instance through the chaos engineering tool, and monitoring the instance through the intelligent monitoring system to determine whether the instance has completed the drill;

[0007] If the drill is completed, sending an alarm message to the drill platform through the intelligent monitoring system, and based on the alarm message, controlling the drill platform to obtain a recovery process in the standard operation process;

[0008] Send the recovery process to the chaos engineering tool through the drill platform. Based on the recovery process, use the chaos engineering tool to recover the instance, and send failure recovery information to the drill platform after the recovery is completed.

[0009] Optionally, the step of, when receiving a failure drill scenario instruction, controlling the drill platform to obtain an instance of a target virtual machine in each of the virtual machines based on the failure drill scenario instruction includes:

[0010] When receiving a failure drill scenario instruction, determine the subsystem corresponding to the failure drill scenario instruction, and determine the virtual machines with instances in each of the virtual machines through the subsystem;

[0011] Control the drill platform to determine a target virtual machine in each of the virtual machines with instances, and obtain the instance of the target virtual machine.

[0012] Optionally, the failure drill system further includes a configuration management module.

[0013] The step of determining the virtual machines with instances in each of the virtual machines through the subsystem includes:

[0014] Obtain subsystem information in the subsystem through the configuration management module, and obtain the virtual machines with instances based on the subsystem information and each of the virtual machines.

[0015] Optionally, the step of, based on the failure drill scenario instruction, using the chaos engineering tool to drill the instance includes:

[0016] Control the chaos engineering tool to determine a failure drill scenario based on the failure drill scenario instruction, establish a failure task through the chaos engineering tool according to the failure drill scenario, and drill the instance through the failure task.

[0017] Optionally, the step of controlling the chaos engineering tool to determine a failure drill scenario based on the failure drill scenario instruction includes:

[0018] Obtain each preset drill scenario in the failure drill system, and control the chaos engineering tool to determine a failure drill scenario in each of the preset drill scenarios through the failure drill scenario instruction;

[0019] If there is no failure drill scenario in each of the preset drill scenarios, obtain the failure drill scenario through the extensible plug-in in the failure drill system.

[0020] Optionally, the step of controlling the drill platform to obtain a recovery process in the standard operation process includes:

[0021] Control the exercise platform to obtain the standard operation process corresponding to the fault exercise scenario in the standard operation process, and obtain a preset recovery instruction, and use the standard operation process and the recovery instruction as the recovery process.

[0022] Optionally, the step of recovering the instance through the chaos engineering tool based on the recovery process and sending a fault recovery message to the exercise platform after the recovery is completed includes:

[0023] Establish a recovery task based on the recovery process, recover the instance according to the recovery task through the chaos engineering tool, and monitor the instance through the intelligent monitoring system to determine whether the instance has been recovered;

[0024] If the instance has been recovered, send a fault recovery message to the exercise platform through the intelligent monitoring system.

[0025] In addition, to achieve the above object, the present invention also provides a fault exercise device, and the fault exercise device includes:

[0026] An acquisition module, configured to, when receiving a fault exercise scenario instruction, control the exercise platform to obtain an instance of a target virtual machine in each virtual machine based on the fault exercise scenario instruction, and send the instance and the fault exercise scenario instruction to the chaos engineering tool through the exercise platform;

[0027] A monitoring module, configured to exercise the instance through the chaos engineering tool based on the fault exercise scenario instruction, and monitor the instance through an intelligent monitoring system to determine whether the instance has been exercised;

[0028] A sending module, configured to, if the exercise is completed, send an alarm message to the exercise platform through the intelligent monitoring system, and based on the alarm message, obtain a recovery process in the standard operation process through the exercise platform;

[0029] A recovery module, configured to send the recovery process to the chaos engineering tool through the exercise platform, recover the instance through the chaos engineering tool based on the recovery process, and send a fault recovery message to the exercise platform after the recovery is completed.

[0030] In addition, to achieve the above object, the present invention also provides a fault exercise device, and the fault exercise device includes: a memory, a processor, and a fault exercise program stored on the memory and executable on the processor, and when the fault exercise program is executed by the processor, the steps of the above-mentioned fault exercise method are implemented.

[0031] In addition, to achieve the above object, the present invention further provides a computer storage medium, on which a fault drill program is stored. When the fault drill program is executed by a processor, the steps of the above-mentioned fault drill method are implemented.

[0032] When the present invention receives a fault drill scenario instruction, it controls the drill platform to obtain an instance of a target virtual machine in each of the virtual machines based on the fault drill scenario instruction, and sends the instance and the fault drill scenario instruction to the chaos engineering tool through the drill platform; based on the fault drill scenario instruction, the chaos engineering tool drills the instance, and the intelligent monitoring system monitors the instance to determine whether the instance has been drilled; if the drilling is completed, the intelligent monitoring system sends an alarm message to the drill platform, and based on the alarm message, the drill platform is controlled to obtain a recovery process in the standard operation process; the drill platform sends the recovery process to the chaos engineering tool, and based on the recovery process, the chaos engineering tool restores the instance and sends a fault recovery message to the drill platform after the restoration is completed. When performing a fault drill, only by obtaining a fault drill scenario instruction can the fault drill start. The whole process from fault generation to fault recovery is completed by the drill platform and related systems without manual participation, that is, a closed-loop drill method is adopted from the start to the end of the fault, with the ability of fault self-healing, thereby reducing the secondary production risk brought by manual operation in the prior art, shortening the fault impact duration, and improving the efficiency of fault drills in the financial system. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 is a schematic diagram of the device structure of the hardware operating environment involved in the embodiment of the present invention;

[0034] Figure 2 is a schematic flowchart of the first embodiment of the fault drill method of the present invention;

[0035] Figure 3 is a schematic diagram of the device modules of the fault drill device of the present invention;

[0036] Figure 4 is a schematic flowchart of the fault drill in the fault drill method of the present invention;

[0037] Figure 5 is a schematic diagram of the fault scenario in the fault drill method of the present invention.

[0038] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] It should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.

[0040] As Figure 1 shown, Figure 1 is a schematic diagram of the device structure of the hardware operating environment involved in the embodiment solution of the present invention.

[0041] The fault drill device in the embodiment of the present invention can be a PC or a server device, on which a Java virtual machine is running.

[0042] As Figure 1 shown, the fault drill device may include: a processor 1001, such as a CPU, a network interface 1004, a user interface 1003, a memory 1005, and a communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0043] Those skilled in the art can understand that Figure 1 the device structure shown in

[0044] does not constitute a limitation on the device, and may include more or fewer components than shown in the figure, or combine certain components, or have a different component layout. Figure 1 shown, the memory 1005, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a fault drill program.

[0045] In Figure 1 the device shown, the network interface 1004 is mainly used to connect to the background server and perform data communication with the background server; the user interface 1003 is mainly used to connect to the client (user side) and perform data communication with the client; and the processor 1001 can be used to call the fault drill program stored in the memory 1005 and execute the operations in the following fault drill method.

[0046] Based on the above hardware structure, an embodiment of the fault drill method of the present invention is proposed.

[0047] Referring to Figure 2 , Figure 2It is a schematic flowchart of the first embodiment of the fault drill method of the present invention, and the method includes:

[0048] Step S10, when receiving a fault drill scenario instruction, control the drill platform to obtain an instance of a target virtual machine in each virtual machine based on the fault drill scenario instruction, and send the instance and the fault drill scenario instruction to the chaos engineering tool through the drill platform;

[0049] It should be noted that in this embodiment, the fault drill method is applied to a fault drill system, and the fault drill system includes: a drill platform, virtual machines (Virtual Machine, VM), a chaos engineering tool, an intelligent monitoring system (Intelligent Monitor System, IMS), and a standard operating procedure (Standard Operating Procedure, SOP). Additionally, in this embodiment, the chaos engineering tool can be replaced by the open-source tool Kube-monkey, and the intelligent monitoring system can be replaced by an open-source operation and maintenance monitoring system.

[0050] In this embodiment, a drill robot can be used to receive, send, and manage terminal instructions (such as fault drill scenario instructions). That is, a fault drill group can be established first, an internal communication group can be established according to the drill scenario, and a drill robot can be installed. The drill communication group can be activated using fixed instructions to display the results of instruction reception and sending. The reception and sending of terminal instructions can be formatted fault scenario terminal instructions, which are pushed to the fault drill platform in the form of parameters using the robot as a bridge to indicate the start of the drill. After the drill ends, the formatted returned data is received and displayed in the fault drill group. The drill platform (which can be a chaos fault drill platform) is built using the Springboot backend framework to construct the chaos fault drill platform, deployed on the production jump server, and dynamically configures the mapping relationship between the drill instructions and operation commands for each scenario to meet the drill requirements of various scenarios, such as hosts (CPU, memory, disk, IO input / output, etc.), networks (packet loss, latency, etc.), DB data (slow queries, table and partition deletion, etc.), and subsystem business transactions (TPS throughput, transaction volume, etc.). And the fault drill platform serves as a data bridge connecting the drill robot and each VM / K8S server and chaos engineering tool, and is docked with the Configuration Management Datebase (CMDB) module, standard operation process, and intelligent monitoring system interface, so that data such as the deployment area of the subsystem to be fault drilled, the IP list of virtual machines and container machines, and the DB can be conveniently obtained. It can also receive the alarm information provided by the intelligent monitoring system in real time, and dock with the standard operation process according to the alarm information. The standard operation process for subsystem anomalies is extracted from the standard operation process and the recovery instructions are issued according to the guidance, realizing a closed-loop drill method. The chaos engineering tool can be obtained by encapsulating the open-source chaos engineering tool using the GO language, and then deployed on the server.

[0051] After the fault drill communication group is established in the terminal, when the drill robot receives the fault drill scenario instruction input by the user, it can connect the subsystem to be fault drilled with the configuration management module, and return all the virtual machine instance IPs of the subsystem through the configuration management module. Then, the fault drill scenario instruction is sent to the drill platform, and the drill platform will select a virtual machine instance IP (such as IP18.192.10.20) as the target virtual machine instance among the virtual machine instance IPs according to the fault drill scenario instruction. Then, the drill platform will send the obtained instance and the fault drill scenario instruction to the chaos engineering tool on the virtual machine server through the SSH (Secure Shell) protocol.

[0052] Step S20: Based on the fault drill scenario instruction, process the instance through the chaos engineering tool, and monitor the instance through the intelligent monitoring system to determine whether the instance fails;

[0053] The intelligent monitoring system can regularly detect the resource status of virtual machines and K8S hosts and collect business transaction information, provide accurate alarm broadcasts for the drill platform, and can also provide parameters for the drill platform to extract the standard process recovery instructions in the standard operation process library. After the drill, the intelligent monitoring system receives the alarm elimination instruction from the drill platform to achieve the function of automatically triggering and eliminating alarms.

[0054] After receiving the fault drill scenario instruction, the chaos engineering tool creates a scenario related to the fault drill scenario instruction for the instance of the target virtual machine. For example, when the fault drill scenario instruction is to select a virtual machine and create an instruction to fill up the CPU for this virtual machine, the chaos engineering tool will create a task to fill up the CPU to process the instance. And in this embodiment, the intelligent monitoring system will monitor the instance of the target virtual machine in real time to determine whether a fault has occurred, that is, to determine whether the chaos engineering tool has completed the processing of the instance. If the processing is completed, it is determined that a fault has occurred.

[0055] Step S30, if a fault occurs, send an alarm message to the drill platform through the intelligent monitoring system. Based on the alarm message, control the drill platform to obtain the recovery process in the standard operation process.

[0056] The standard operation process configures the standard operation processes for the fault scenarios of each subsystem host, DB, network, business, etc. The fault drill platform matches the corresponding standard operation process for recovering the fault in the standard operation process according to the alarm information synchronized by the intelligent monitoring system, and sends the standard operation process to the fault drill group and the chaos engineering tool. When it is determined that a fault has occurred in the instance of the target virtual machine, that is, when the chaos engineering tool has completed the processing of the instance, an alarm message can be sent to the drill platform through the intelligent monitoring system. It should be noted that the intelligent monitoring system will also send an alarm message to the drill robot while sending the alarm message to the drill platform. After receiving the alarm message, the drill platform will search the standard operation process library according to the content in the alarm message and the subsystem to match the corresponding fault recovery standard operation process, that is, the recovery process.

[0057] Step S40, send the recovery process to the chaos engineering tool through the drill platform. Based on the recovery process, use the chaos engineering tool to recover the instance, and send a fault recovery message to the drill platform after the recovery is completed.

[0058] After obtaining the recovery process through the rehearsal platform, the recovery process will be sorted out to form a recovery instruction, and this recovery instruction will be sent to the chaos engineering tool. The chaos engineering tool will establish a corresponding recovery task according to the recovery instruction and perform a recovery operation on the faulty instance through the recovery task. For example, a recovery operation is performed on the CPU full-load fault of the target virtual machine instance until the recovery task is completed. Since the intelligent monitoring system monitors the target virtual machine instance in real time, after the intelligent monitoring system detects that the target virtual machine instance has returned to normal, it will turn off the intelligent monitoring system alarm and push the fault recovery information to the rehearsal platform. The rehearsal platform will announce the end of the rehearsal and format the summary data, and send the rehearsal summary to the terminal fault rehearsal communication group through the rehearsal robot.

[0059] In addition, to assist in understanding the working principle of the closed-loop fault rehearsal, an example is given below.

[0060] For example, as Figure 4As shown in the figure, the closed-loop fault drill mainly consists of a drill robot, a chaotic fault drill platform, virtual machines, and K8S (Kubernetes Google) container servers, a chaos engineering tool, an intelligent monitoring system, a configuration management module, a standard operation process function system, and a fault drill scenario. The fault drill scenario can include host CPU, memory, IO (input / output), network packet loss, latency, network disconnection, DB (data) slow query, table partitioning, master-slave delay, subsystem TPS (throughput), transaction volume, and process killing, etc. Among them, the chaotic fault drill platform is also equipped with an extensible plug-in and a drill configuration management terminal. And the number of virtual machines is the same as the number of containers in K8S. That is, assuming there are virtual machines 1, virtual machines 2, virtual machines 3, virtual machines 4, etc., then there are also containers 1, containers 2, containers 3, containers 4, etc. in K8S. The chaotic fault drill process can be that the terminal issues a fault drill scenario instruction to the drill robot, and the subsystem exports a virtual machine instance with subsystem information through the configuration management module. The chaotic fault drill platform randomly selects a virtual machine instance according to the fault drill scenario instruction, and sends the obtained instance and the fault drill scenario instruction to the chaos engineering tool on the virtual machine server. The chaos engineering tool processes the obtained instance according to the fault drill scenario instruction. After the intelligent monitoring system monitors that the chaos engineering tool has completed the processing of the instance, it sends an alarm message to the chaotic fault drill platform and the drill robot. The chaotic fault drill platform obtains the standard operation process for fault recovery in the standard operation process according to the alarm message and passes it to the chaos engineering tool. The chaos engineering tool restores the instance through the chaos engineering tool. After the intelligent monitoring system monitors that the chaos engineering tool has restored the instance and the restoration is completed, it turns off the alarm of the intelligent monitoring system and pushes the alarm message to the fault drill platform. At this time, the closed-loop fault drill has been completed. Among them, the extensible plug-in is used to supplement the fault drill scenario when it is found that there is no fault drill scenario required by the user during the closed-loop fault drill.

[0061] In this embodiment, when a fault drill scenario instruction is received, an instance of the target virtual machine is obtained in each of the virtual machines based on the fault drill scenario instruction, and the instance and the fault drill scenario instruction are sent to the chaos engineering tool through the drill platform; based on the fault drill scenario instruction, the instance is drilled through the chaos engineering tool, and the instance is monitored through the intelligent monitoring system to determine whether the instance has completed the drill; if the drill is completed, an alarm message is sent to the drill platform through the intelligent monitoring system, and based on the alarm message, a recovery process is obtained in the standard operation process through the drill platform; the recovery process is sent to the chaos engineering tool through the drill platform, and based on the recovery process, the instance is recovered through the chaos engineering tool, and a fault recovery message is sent to the drill platform after the recovery is completed. When performing a fault drill, only by obtaining the fault drill scenario instruction can the fault drill start. The entire process from fault generation to fault recovery is completed by the drill platform and associated systems without manual intervention, that is, a closed-loop drill method is adopted from the start to the end of the fault, with the ability of fault self-healing, thus reducing the secondary production risk brought by manual operation in the prior art, shortening the fault impact duration, and improving the efficiency of fault drills in the financial system.

[0062] Further, based on the first embodiment of the fault drill method of the present invention, a second embodiment of the fault drill method of the present invention is proposed. This embodiment is a refinement of step S10 in the first embodiment of the present invention, which is to control the drill platform to obtain an instance of the target virtual machine in each of the virtual machines based on the fault drill scenario instruction when the fault drill scenario instruction is received, and includes:

[0063] Step a, when the fault drill scenario instruction is received, determine the subsystem corresponding to the fault scenario instruction, and determine the virtual machines with instances in each of the virtual machines through the subsystem;

[0064] The subsystem can be the system that needs to perform the fault drill. When the fault drill scenario instruction is received, first determine the subsystem that needs to perform the fault drill through the fault drill scenario instruction, and pass the subsystem into the configuration management module. The instances of all virtual machines with subsystem information are exported through the configuration management module.

[0065] Step b, control the drill platform to determine the target virtual machine in each of the virtual machines with instances, and obtain the instance of the target virtual machine.

[0066] After obtaining each virtual machine with an instance, the drill platform will select a virtual machine that meets the requirements of the fault drill scenario instruction in each virtual machine as the target virtual machine according to the fault drill scenario instruction, and obtain the instance of the target virtual machine.

[0067] In this embodiment, when a fault drill scenario instruction is received, the subsystem is determined, and the virtual machines with instances are determined through the subsystem, so as to further obtain the instances of the target virtual machines, ensuring the accuracy of the instances of the target virtual machines obtained during the fault drill.

[0068] Specifically, the step of determining the virtual machines with instances in each of the virtual machines through the subsystem includes:

[0069] Step a1, obtain the subsystem information of the subsystem through the configuration management module, and obtain the virtual machines with instances based on the subsystem information and each of the virtual machines.

[0070] It should be noted that in this embodiment, the fault drill system further includes a configuration management module.

[0071] The configuration management module configures information such as the deployment areas of each subsystem, virtual machines and container instances, subsystem DB, and DB instances, providing target data for the fault drill platform. After the subsystem is determined, the subsystem information of the subsystem (such as the subsystem DB, etc.) can be obtained through the configuration management module, and the virtual machines with instances can be established through the subsystem information and each controlled virtual machine. Among them, there can be multiple virtual machines with instances.

[0072] In this embodiment, the configuration management module is used to obtain the virtual machines with instances associated with the subsystem, ensuring that the obtained virtual machines with instances are all associated with the subsystem and ensuring the accuracy of the fault drill.

[0073] Further, the step of processing the instance through the chaos engineering tool based on the fault drill scenario instruction includes:

[0074] Step c, control the chaos engineering tool to determine a preset fault drill scenario based on the fault drill scenario instruction, establish a fault task through the chaos engineering tool according to the preset fault drill scenario, and conduct a drill on the instance through the fault task.

[0075] When the chaos engineering tool obtains the fault drill scenario instruction, it will determine the preset fault drill scenario that needs to be carried out for the fault drill according to the fault drill scenario instruction, establish a fault task according to the preset fault drill scenario, and perform fault processing (i.e., drill) on the obtained instance according to the fault task to establish a fault scenario. When the fault processing is completed, the fault scenario is constructed at this time, and the intelligent monitoring system alarm system will timely monitor this fault scenario and give an alarm prompt.

[0076] In this embodiment, a fault task is established according to a fault drill scenario by a chaos engineering tool to drill the instance, thereby ensuring the accuracy of the fault drill.

[0077] Specifically, the steps of controlling the chaos engineering tool to determine the fault drill scenario based on the fault drill scenario instruction include:

[0078] Step c1, obtain each preset drill scenario in the fault drill system, and control the chaos engineering tool to determine the fault drill scenario from each of the preset drill scenarios through the fault drill scenario instruction;

[0079] Obtain each preset drill scenario (i.e., each preset drill scenario) in the fault drill system, and when the fault drill scenario instruction is obtained, screen out the drill scenario that matches the fault drill scenario instruction from each drill scenario, and use it as the fault drill scenario. Among them, the drill scenarios cover hosts, networks, databases, subsystem transactions, etc., and the mapping relationship between the drill instruction and the specific operation command needs to be configured in the drill configuration management end of the fault drill system in advance. The settings of the drill scenarios can be customized according to the drill requirements. For example, as Figure 5 shown, some drill scenarios include applications and data, (kill process, suspend process, heartbeat exception, startup exception, environment error, packet error or damage, configuration misdeletion or error or acquisition exception, system single point, dependency exception, dependency timeout, asynchronous blocking synchronization, memory overflow, thread pool full, monitoring error, flow control unreasonable). During operation, the middleware and the system can also include (load balancing failure, cache current limiting, cache hotspot, database hotspot, database downtime, data synchronization delay, table partition misdeletion, database connection full, data master-slave delay, memory preemption, memory error, context switching, CPU preemption). Another example is during operation, in the middleware, the system, and the system (server downtime, server false death, power off, unwritable, unreadable, disk full or bad or slow, co-location, network card full, network packet loss or jitter, network timeout, network disconnection, domain name system failure, network timeout). System services (system throughput sudden increase, transaction volume sudden increase, transaction latency, transaction blockage).

[0080] Step c2, if there is no fault drill scenario in each of the preset drill scenarios, obtain the fault drill scenario through the extensible plug-in in the fault drill system.

[0081] When it is determined through judgment that there is no fault drill scenario in each drill scenario, the fault drill scenario can be obtained through the extensible plug-in in the fault drill system. Among them, the extensible plug-in is a plug-in developed using the Java language, which mainly provides business fault drill scenarios with configurable modes such as subsystems and transaction interfaces, has the ability of multi-threaded high concurrency, can call the subsystem transaction interface based on the HTTP protocol, create a fault scenario from the direction of business traffic, and the application fault drill scenario mainly focuses on the subsystem business transaction scope, such as transaction volume, transaction latency, etc.

[0082] In this embodiment, by determining the fault drill scenario in each drill scenario, when it is found that there is no fault drill scenario in each drill scenario, the fault drill scenario is obtained through the extensible plug-in, thereby improving the drill scope of the fault drill method.

[0083] Further, the steps of controlling the drill platform to obtain the recovery process in the standard operation process include:

[0084] Step d, controlling the drill platform to obtain the standard operation process corresponding to the fault drill scenario in the standard operation process, and obtaining a preset recovery instruction, and taking the standard operation process and the recovery instruction as the recovery process.

[0085] When the drill platform receives an alarm message, the drill platform will determine the standard operation process corresponding to the fault drill scenario in the standard operation process, and will sort it out in the drill platform, and then take the recovery instruction and the standard operation process as the recovery process and transfer it to the chaos engineering tool to recover the instance of the target virtual machine.

[0086] In this embodiment, by obtaining the standard operation process and the recovery instruction and taking them as the recovery process, the chaos engineering tool can perform the recovery operation normally, ensuring the normal progress of the fault drill method.

[0087] Further, the steps of recovering the instance through the chaos engineering tool based on the recovery process and sending a fault recovery message to the drill platform after the recovery is completed include:

[0088] Step f, establishing a recovery task based on the recovery process, recovering the instance according to the recovery task through the chaos engineering tool, and monitoring the instance through the intelligent monitoring system to determine whether the instance has been recovered;

[0089] When the chaos engineering tool obtains the recovery process, it can establish recovery tasks (such as automatic elimination, isolation and restart, and automatic scaling) according to the recovery process, and perform recovery operations on the instances of the target virtual machine according to the recovery tasks until the recovery is completed. Since the intelligent monitoring system will monitor the status of the instances of the target virtual machine in real time, the intelligent monitoring system can be used to determine whether the target virtual machine has been recovered.

[0090] Step g, if the instance has been recovered, the intelligent monitoring system is used to send a fault recovery message to the drill platform.

[0091] When it is determined that the instance of the target virtual machine has been recovered, a fault recovery message can be sent to the drill platform through the intelligent monitoring system. At this time, the drill platform will output a prompt message indicating the end of the drill.

[0092] In this embodiment, when the chaos engineering tool recovers the instance, the intelligent monitoring system is used for monitoring, and when the instance is recovered, a fault recovery message is sent to the drill platform, thus ensuring the efficiency of the fault drill.

[0093] The present invention also provides a fault drill device. Referring to Figure 3 , the fault drill device includes:

[0094] An acquisition module, configured to, when receiving a fault drill scenario instruction, control the drill platform to obtain an instance of a target virtual machine in each virtual machine based on the fault drill scenario instruction, and send the instance and the fault drill scenario instruction to the chaos engineering tool through the drill platform;

[0095] A monitoring module, configured to, based on the fault drill scenario instruction, perform a drill on the instance through the chaos engineering tool, and monitor the instance through the intelligent monitoring system to determine whether the instance has completed the drill;

[0096] A sending module, configured to, if the drill is completed, send an alarm message to the drill platform through the intelligent monitoring system, and based on the alarm message, obtain a recovery process in the standard operation process through the drill platform;

[0097] A recovery module, configured to send the recovery process to the chaos engineering tool through the drill platform, perform recovery on the instance through the chaos engineering tool based on the recovery process, and send a fault recovery message to the drill platform after the recovery is completed.

[0098] Optionally, the acquisition module is further configured to:

[0099] When receiving a fault drill scenario instruction, determine the subsystem corresponding to the fault scenario instruction, and determine the virtual machines with instances among all the virtual machines through the subsystem;

[0100] Control the drill platform to determine target virtual machines among all the virtual machines with instances, and obtain the instances of the target virtual machines.

[0101] Optionally, the fault drill system further includes a configuration management module, and the obtaining module is further configured to:

[0102] Obtain subsystem information in the subsystem through the configuration management module, and determine the virtual machines with instances based on the subsystem information and all the virtual machines.

[0103] Optionally, the monitoring module is further configured to:

[0104] Control the chaos engineering tool to determine a fault drill scenario based on the fault drill scenario instruction, establish a fault task according to the fault drill scenario through the chaos engineering tool, and conduct a drill on the instance through the fault task.

[0105] Optionally, the monitoring module is further configured to:

[0106] Obtain each preset drill scenario in the fault drill system, and control the chaos engineering tool to determine a fault drill scenario in each preset drill scenario through the fault drill scenario instruction;

[0107] If there is no fault drill scenario in each preset drill scenario, obtain the fault drill scenario through the extensible plugin in the fault drill system.

[0108] Optionally, the monitoring module is further configured to:

[0109] Control the drill platform to obtain the standard operation process corresponding to the fault drill scenario in the standard operation process, and obtain a preset recovery instruction, and use the standard operation process and the recovery instruction as the recovery process.

[0110] Optionally, the recovery module is further configured to:

[0111] Establish a recovery task based on the recovery process, recover the instance according to the recovery task through the chaos engineering tool, and monitor the instance through the intelligent monitoring system to determine whether the instance has been recovered;

[0112] If the instance has been recovered, send a fault recovery message to the drill platform through the intelligent monitoring system.

[0113] The methods executed by the above program modules can refer to the various embodiments of the fault drill method of the present invention, which will not be elaborated here.

[0114] The present invention also provides a computer storage medium.

[0115] A fault drill program is stored on the computer storage medium of the present invention. When the fault drill program is executed by a processor, the steps of the fault drill method described above are implemented.

[0116] Among them, the method implemented when the fault drill program running on the processor is executed can refer to the various embodiments of the fault drill method of the present invention, which will not be elaborated here.

[0117] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such a process, method, article or system. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article or system including that element.

[0118] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0119] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0120] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied to other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A fault drill method, characterized in that, the fault drill method is applied to a fault drill system, and the fault drill system includes: a drill platform, multiple virtual machines, a chaos engineering tool, an intelligent monitoring system, and a standard operation process, the fault drill method includes the following steps: When receiving a fault drill scenario instruction, controlling the drill platform to obtain an instance of a target virtual machine in each of the virtual machines based on the fault drill scenario instruction, and sending the instance and the fault drill scenario instruction to the chaos engineering tool through the drill platform; Based on the fault drill scenario instruction, using the chaos engineering tool to drill the instance, and using the intelligent monitoring system to monitor the instance to determine whether the instance has completed the drill; If the drill is completed, sending an alarm message to the drill platform through the intelligent monitoring system, and based on the alarm message, controlling the drill platform to obtain a recovery process in the standard operation process; Sending the recovery process to the chaos engineering tool through the drill platform, and based on the recovery process, using the chaos engineering tool to recover the instance, and sending a fault recovery message to the drill platform after the recovery is completed; The step of using the chaos engineering tool to drill the instance based on the fault drill scenario instruction includes: Controlling the chaos engineering tool to determine a fault drill scenario based on the fault drill scenario instruction, establishing a fault task according to the fault drill scenario through the chaos engineering tool, and using the fault task to drill the instance; The step of controlling the chaos engineering tool to determine a fault drill scenario based on the fault drill scenario instruction includes: Obtaining each preset drill scenario in the fault drill system, and controlling the chaos engineering tool to determine a fault drill scenario in each of the preset drill scenarios through the fault drill scenario instruction; If there is no fault drill scenario in each of the preset drill scenarios, obtaining the fault drill scenario through an extensible plug-in in the fault drill system.

2. The fault drill method according to claim 1, characterized in that, the step of, when receiving a fault drill scenario instruction, controlling the drill platform to obtain an instance of a target virtual machine in each of the virtual machines based on the fault drill scenario instruction includes: When receiving a fault drill scenario instruction, determining the subsystem corresponding to the fault drill scenario instruction, and determining the virtual machines with instances in each of the virtual machines through the subsystem; Controlling the drill platform to determine a target virtual machine in each of the virtual machines with instances, and obtaining an instance of the target virtual machine.

3. The fault drill method according to claim 2, characterized in that, the fault drill system further includes a configuration management module, the step of determining the virtual machines with instances in each of the virtual machines through the subsystem includes: Obtaining subsystem information in the subsystem through the configuration management module, and obtaining the virtual machines with instances based on the subsystem information and each of the virtual machines.

4. The fault drill method according to claim 1, It is characterized in that The steps of controlling the drill platform to obtain the recovery process in the standard operation process include: Controlling the drill platform to obtain the standard operation process corresponding to the fault drill scenario in the standard operation process, obtaining a preset recovery instruction, and using the standard operation process and the recovery instruction as the recovery process.

5. The fault drill method according to any one of claims 1-4, It is characterized in that The steps of recovering the instance through the chaos engineering tool based on the recovery process and sending fault recovery information to the drill platform after the recovery is completed include: Establishing a recovery task based on the recovery process, recovering the instance according to the recovery task through the chaos engineering tool, and monitoring the instance through the intelligent monitoring system to determine whether the instance has been recovered; If the instance has been recovered, the intelligent monitoring system is used to send fault recovery information to the drill platform.

6. A fault drill device, It is characterized in that The fault drill device includes: An acquisition module, configured to, when receiving a fault drill scenario instruction, control the drill platform to obtain an instance of a target virtual machine in each virtual machine based on the fault drill scenario instruction, and send the instance and the fault drill scenario instruction to the chaos engineering tool through the drill platform; A monitoring module, configured to drill the instance through the chaos engineering tool based on the fault drill scenario instruction, and monitor the instance through the intelligent monitoring system to determine whether the instance has been drilled; The monitoring module is further configured to: Control the chaos engineering tool to determine a fault drill scenario based on the fault drill scenario instruction, establish a fault task according to the fault drill scenario through the chaos engineering tool, and drill the instance through the fault task; The monitoring module is further configured to: Obtain each preset drill scenario in the fault drill system, and control the chaos engineering tool to determine a fault drill scenario in each of the preset drill scenarios through the fault drill scenario instruction; If there is no fault drill scenario in each of the preset drill scenarios, obtain the fault drill scenario through an extensible plug-in in the fault drill system; A sending module, configured to, if the drill is completed, send an alarm message to the drill platform through the intelligent monitoring system, and based on the alarm message, obtain a recovery process in the standard operation process through the drill platform; A recovery module, configured to send the recovery process to the chaos engineering tool through the drill platform, recover the instance through the chaos engineering tool based on the recovery process, and send fault recovery information to the drill platform after the recovery is completed.

7. A fault drill device, It is characterized in that The fault drill device includes: a memory, a processor, and a fault drill program stored on the memory and executable on the processor. When the fault drill program is executed by the processor, the steps of the fault drill method according to any one of claims 1 to 5 are implemented.

8. A computer storage medium, It is characterized in that a fault drill program is stored on the computer storage medium, and when the fault drill program is executed by a processor, the steps of the fault drill method described in any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Method for fault-injection test based on virtual machine

    CN101872323A

  • Power distribution network fault section locating method based on adaptive chaotic fruit fly optimization algorithm

    CN107064731A