Fault drilling method and device, electronic equipment and storage medium

Through automated fault drill methods and an orchestration strategy based on a set of fault scenario cases, the low efficiency and low accuracy of manual fault injection testing are solved, efficient and accurate fault testing is achieved, and the stability and security of the system are improved.

CN120705029APending Publication Date: 2025-09-26NETSUNION CLEARING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410339239.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-22
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

When the existing technology performs testing by manually injecting faults, the workload is large and errors are prone to occur, resulting in low accuracy and execution efficiency of fault testing.

Method used

Based on the fault scenario case set, target fault scenario cases are obtained, the drill sequence is determined, and the drills are arranged through preset orchestration strategies to generate a drill plan, thereby realizing automated fault drills and reducing manual participation.

Benefits of technology

It improves the accuracy and execution efficiency of fault testing, can more accurately identify system weaknesses, and enhance system stability and security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705029A_ABST
    Figure CN120705029A_ABST
Patent Text Reader

Abstract

The invention discloses a fault drilling method and device, electronic equipment and a storage medium, and the method comprises the steps: obtaining N target fault scene cases from a fault scene case set based on a fault drilling demand for a business system, and determining a drilling sequence among the N target fault scene cases; based on a preset arrangement strategy and the drilling sequence, arranging the N target fault scene cases to obtain M groups of drilling plans, each group of drilling plans including at least one fault scene case, and the preset arrangement strategy including at least one of the following arrangement modes: parallel arrangement and serial arrangement; and performing fault drilling on the business system based on the M groups of drilling plans.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computers, and in particular to a fault drill method, device, electronic device, and storage medium. Background Art

[0002] As businesses continue to grow and their system architectures become increasingly complex, maintaining business continuity remains paramount. While traditional quality assurance methods are essential to ensure business continuity, proactively identifying system weaknesses through external fault injection is becoming increasingly essential. Currently, testing system performance under fault conditions is primarily done manually through fault injection. This involves a series of steps, including preparing the environment, preparing fault scenarios, manually injecting faults, observing system performance, and restoring the environment after testing is complete.

[0003] However, the method of testing by manually injecting faults is labor-intensive and error-prone, requiring a large amount of time and manpower, resulting in problems such as low fault testing accuracy and reduced execution efficiency. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to provide a fault drill method, device, electronic device and storage medium to solve the problem of reduced test accuracy and execution efficiency caused by manual fault testing.

[0005] In order to achieve the above objectives, the embodiments of the present application adopt the following technical solutions:

[0006] In a first aspect, an embodiment of the present application provides a fault drill method, comprising:

[0007] Based on the fault drill requirements for the business system, N target fault scenario cases are obtained from the fault scenario case set, and a drill order between the N target fault scenario cases is determined;

[0008] Based on a preset orchestration strategy and the drill sequence, the N target fault scenario cases are orchestrated to obtain M groups of drill plans, each group of drill plans including at least one fault scenario case, the preset orchestration strategy including at least one of the following orchestration modes: parallel orchestration and serial orchestration;

[0009] Based on the M group drill plan, a fault drill is performed on the business system.

[0010] In a second aspect, an embodiment of the present application provides a fault drill device, comprising:

[0011] An acquisition unit is configured to acquire N target fault scenario cases from a fault scenario case set based on a fault drill requirement for the business system, and determine a drill order among the N target fault scenario cases;

[0012] An orchestration unit is configured to orchestrate the N target fault scenario cases based on a preset orchestration strategy and the drill sequence to obtain M groups of drill plans, each group of drill plans including at least one fault scenario case, wherein the preset orchestration strategy includes at least one of the following orchestration modes: parallel orchestration and serial orchestration;

[0013] A drill unit is used to perform a fault drill on the business system based on the M groups of drill plans.

[0014] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a processor; and a memory for storing instructions executable by the processor; wherein the processor is configured to execute the instructions to implement the method described in the first aspect.

[0015] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which, when instructions in the storage medium are executed by a processor of an electronic device, enables the electronic device to execute the method described in the first aspect.

[0016] At least one of the above-mentioned technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects: based on the fault drill requirements for the business system, N target fault scenario cases are obtained from the fault scenario case set, and the drill order between the N target fault scenario cases is determined; based on the preset orchestration strategy and the drill order, the N target fault scenario cases are orchestrated to obtain M groups of drill plans, which meet the requirements of complex scenarios for fault drills of the business system, and make the organization of fault scenario cases highly liberalized without the need for manual preparation; based on the M groups of drill plans, a fault drill is performed on the business system, and after the entire automated drill process is started, manual participation can be completely eliminated, and multimodal intelligent control and automatic drills of fault scenario cases can be realized, thereby saving a lot of manpower and time, improving execution efficiency and test accuracy, and being able to more accurately discover weaknesses in the system for better upgrade and maintenance, and enhancing the stability and security of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0018] Figure 1 A flowchart of a fault drill method provided in one embodiment of the present application;

[0019] Figure 2 A schematic diagram of a structure for arranging fault scenario cases provided in one embodiment of the present application;

[0020] Figure 3A One of the flowcharts of another fault drill method provided in one embodiment of the present application;

[0021] Figure 3B A second flowchart of another fault drill method provided in one embodiment of the present application;

[0022] Figure 4 A schematic structural diagram of a fault drill device provided in one embodiment of the present application;

[0023] Figure 5 A schematic structural diagram of an electronic device provided in accordance with an embodiment of the present application. DETAILED DESCRIPTION

[0024] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0025] The terms "first," "second," and the like in this specification and claims are used to distinguish similar objects and are not intended to describe a particular order or precedence. It should be understood that such terms are interchangeable, where appropriate, so that the embodiments of the present application may be implemented in an order other than that illustrated or described herein. Furthermore, the term "and / or" in this specification and claims refers to at least one of the connected objects, and the character " / " generally indicates that the connected objects are in an "or" relationship.

[0026] The following describes in detail the technical solutions provided by various embodiments of the present application in conjunction with the accompanying drawings.

[0027] The fault drill method provided in the embodiment of the present application is applicable to various scenarios where fault testing of the system is required. With the continuous development of business, the business system architecture is becoming more and more complex, and maintaining business continuity is always a top priority. While relying on traditional quality assurance methods to ensure business continuity, the need to actively discover the weaknesses of the system through the injection of external faults is becoming more and more indispensable. At present, the performance of the system under faults is mainly tested by injecting faults manually, including a series of operations such as environment preparation, preparation of fault scenario cases, manual injection of faults, observation of system performance, and restoration of the environment after the test is completed.

[0028] However, the testing method of manually injecting faults is complicated, labor-intensive, and error-prone, requiring a large amount of time and manpower, resulting in low accuracy of fault testing and reduced execution efficiency.

[0029] In view of this, an embodiment of the present application provides a fault drill method, which, based on the fault drill requirements for the business system, obtains N target fault scenario cases from a fault scenario case set, and determines the drill order between the N target fault scenario cases. Based on a preset orchestration strategy and the drill order, the N target fault scenario cases are orchestrated to obtain M groups of drill plans, which meet the requirements of complex scenarios for fault drills of the business system, and make the organization of the fault scenario cases highly liberalized without the need for manual preparation. Based on the M groups of drill plans, a fault drill is performed on the business system. After the entire automated drill process is started, manual participation can be completely eliminated, and multi-modal intelligent control and automatic drills of the fault scenario cases can be realized, thereby saving a lot of manpower and time, improving execution efficiency and test accuracy, and being able to more accurately discover weaknesses in the system for better upgrades and maintenance, and enhancing the stability and security of the system.

[0030] Specifically, see Figure 1 , is a flowchart of a fault drill method provided by an embodiment of the present application. Figure 1 As shown, the fault drill method provided in the embodiment of the present application includes the following steps:

[0031] S102 , based on the fault drill requirements for the business system, obtain N target fault scenario cases from the fault scenario case set, and determine a drill sequence among the N target fault scenario cases.

[0032] Among them, fault drill requirements refer to a series of requirements or standards formulated based on test objectives and system characteristics before conducting a fault drill. N is a positive integer. The drill sequence refers to the order in which fault scenario cases are injected into the business system for drills. The target fault scenario case refers to the fault scenario case required for fault drills for the business system. The fault scenario case set is a collection of several fault scenario cases. Fault scenario cases are composed of fault case parameters. Fault case parameters refer to parameters that describe the specific characteristics, conditions or attributes of the fault scenario. A fault scenario refers to the scenario when a business system fails. By introducing different fault case parameters, the state, behavior or configuration of the business system under a specific fault scenario can be simulated to discover the weaknesses of the business system, and then the business system can be upgraded based on the weaknesses to ensure the robustness and reliability of the business system under various abnormal conditions.

[0033] It should be understood that each fault scenario case corresponds to a fault scenario. As another optional implementation, when each target fault scenario case corresponds to a fault scenario, the following steps are further included before the above S102: based on the fault drill requirements, determine N fault scenarios corresponding to the business system;

[0034] For each fault scenario, split the fault scenario into multiple atomic scenarios according to the scenario category;

[0035] Obtain the fault case parameters corresponding to each atomic scenario from the atomic fault case set;

[0036] The fault case parameters of multiple atomic scenarios contained in the fault scenario are combined to obtain the target fault scenario case corresponding to the fault scenario and add it to the fault scenario case set.

[0037] Scenario categories refer to fault scenario classifications or types. These classifications can be made based on the nature and impact of faults, facilitating the organization and planning of fault drills. Atomic scenarios are the fundamental units for fault drills in business systems. They describe the smallest, most basic fault scenarios for specific components, modules, or functions in a business system under certain abnormal conditions. An atomic fault case set consists of several fault case parameters.

[0038] According to the fault drill requirements, determine the N fault scenarios to be simulated for the business system drill. According to the type of each fault scenario, split each fault scenario into multiple atomic scenarios. Then obtain the fault case parameters required to simulate these atomic scenarios from the atomic fault case set. Combine these fault case parameters to obtain the target fault scenario case corresponding to the fault scenario and add it to the fault scenario case set.

[0039] The embodiment of the present application splits the fault scenario into atomic scenarios, which can more finely control the occurrence conditions and impact scope of each fault, more accurately simulate specific fault scenarios, make the test modular, and can be tested individually, so as to locate and identify system problems and potential weaknesses, which helps to troubleshoot and repair faults more quickly. Conversely, by combining different atomic scenarios to obtain larger fault scenarios, more complex and diverse fault scenarios can be created to ensure the robustness of the system in the face of multiple faults.

[0040] S104: Based on the preset arrangement strategy and drill sequence, N target fault scenario cases are arranged to obtain M groups of drill plans.

[0041] Where M is a positive integer. Each drill plan includes at least one fault scenario case. The preset orchestration strategy includes at least one of the following orchestration methods: parallel orchestration and serial orchestration.

[0042] Parallel orchestration refers to the simultaneous execution of multiple target failure scenarios during a fault drill. For example, if the drill order for Failure Scenario Case A, Failure Scenario Case B, and Failure Scenario Case C is the same, then these scenarios need to be run simultaneously. After parallel orchestration, the fault drill should execute Failure Scenario Case A, Failure Scenario Case B, and Failure Scenario Case C simultaneously.

[0043] Serial orchestration refers to the arrangement of executing target failure scenarios one by one in a specific order during a fault drill. For example, if you orchestrate failure scenarios A, B, and C in parallel, the drill order would be Failure Scenario A followed by Failure Scenario C, which in turn would be followed by Failure Scenario B. When performing the fault drill after serial orchestration, Failure Scenario B should be executed first, followed by Failure Scenario C, and finally by Failure Scenario A.

[0044] As an optional implementation, the above S104 may include the following steps: based on the drill order between the N target fault scenario cases, arranging the target fault scenario cases to be drilled simultaneously into the same group to obtain T case groups;

[0045] T case groups are serially arranged to obtain M groups of exercise plans.

[0046] Here, T is a positive integer. The above arrangement process will be further illustrated with a specific example, which should not be understood as limiting the method of the embodiment of the present application. Example: According to the fault drill requirements, fault scenario case a, fault scenario case b, fault scenario case c, and fault scenario case d are obtained from the fault scenario case set as target fault scenario cases, that is, N = 4. Among them, fault scenario case a and fault scenario case b are drilled simultaneously, and are arranged in parallel to obtain a corresponding case group, recorded as case group A. Fault scenario case a, fault scenario case b, and fault scenario case c are drilled simultaneously, and are arranged in parallel to obtain a corresponding case group, recorded as case group B. Fault scenario case a and fault scenario case d are drilled simultaneously, and are arranged in parallel to obtain a corresponding case group, recorded as case group C. There are a total of 3 case groups, that is, T = 3. These 3 case groups are serially arranged, and case group A and case group B are serially arranged to obtain a corresponding set of drill plans, recorded as drill plan I. Case group A and case group B are serially arranged to obtain a corresponding set of drill plans, recorded as drill plan II. There are a total of 2 sets of drill plans, that is, M = 2.

[0047] It should be understood that the embodiment of the present application is to first perform concurrent orchestration on the target fault scenario cases to obtain multiple case groups, and then perform serial orchestration on the multiple case groups to obtain each group of exercise plans. Alternatively, the target fault scenario cases may be first performed serially to obtain multiple case groups, and then perform parallel orchestration on the multiple cases to obtain each group of exercise plans. For example: assuming that fault scenario case a is executed before fault scenario case b, case group A is obtained through serial orchestration. Fault scenario case c is performed after fault scenario case a and fault scenario case b, and case group B is obtained through serial orchestration. Fault scenario case c is performed before fault scenario case d, and case group C is obtained through serial orchestration. Case group A, case group B, and case group C are then orchestrated in parallel. Each case group corresponds to a set of exercise plans, i.e., M=3. When performing a fault drill, these three sets of exercise plans are executed simultaneously.

[0048] The embodiment of the present application arranges N target fault scenario cases based on a preset orchestration strategy and drill sequence to obtain M groups of drill plans, organically integrating the fault scenario cases to meet the execution requirements of the business system for fault injection, making the organization of the fault scenario cases highly liberalized, eliminating the need for manual preparation of fault scenario cases, and thus realizing automated fault drills for the business system, saving manpower and time costs, and improving the accuracy and execution efficiency of fault testing.

[0049] S106: Based on the drill plan of group M, a fault drill is performed on the business system.

[0050] The embodiment of the present application is based on the fault drill requirements for the business system, obtains N target fault scenario cases from a fault scenario case set, and determines the drill order between the N target fault scenario cases. Based on a preset orchestration strategy and the drill order, the N target fault scenario cases are orchestrated to obtain M groups of drill plans, which meet the requirements of complex scenarios for fault drills of the business system, and make the organization of the fault scenario cases highly liberalized without the need for manual preparation. Based on the M groups of drill plans, a fault drill is performed on the business system. After the entire automated drill process is started, manual participation can be completely eliminated, and multi-modal intelligent control and automatic drills of the fault scenario cases can be achieved, thereby saving a lot of manpower and time, improving execution efficiency and test accuracy, and being able to more accurately discover weaknesses in the system for better upgrades and maintenance, and enhancing the stability and security of the system.

[0051] As an optional implementation, when each group of drill plans includes multiple case groups arranged in series, and each case group includes at least one target failure scenario case, the above S106 includes the following steps:

[0052] S160: Determine an execution strategy for the currently acquired drill plan based on the fault drill requirement.

[0053] S162: Add M groups of drill plans to the cache queue.

[0054] S164: Consume the cache queue and perform a fault drill on the business system based on the execution strategy of the currently acquired drill plan and the case groups included in each group of drill plans.

[0055] The execution strategy includes at least one of the following execution modes: sequential execution and random execution. Based on the fault drill requirements, the execution strategy corresponding to each group of drill plans is determined. It should be understood that the execution strategies corresponding to each of the M groups of drill plans can be different to meet the diverse needs of intelligently controlling fault scenario cases in business system fault drills.

[0056] Among them, the cache queue generally refers to a data structure used for temporary storage of data. M groups of drill plans are added to the cache queue. Each time a group of drill plans is consumed, the group of drill plans is locked to ensure that only one thread can perform fault drills on the business system based on the group of drill plans at the same time, so as to avoid the problem of data inconsistency caused by multiple threads operating the same group of drill plans concurrently, and then unlock it to continue consuming the next group of drill plans. In actual applications, the local cache queue can be implemented using various message middleware, such as kafka, redis and other commonly used message middleware in this field, or independently developed message middleware can be used, etc., and the embodiments of this specification do not limit this.

[0057] After determining the execution strategy corresponding to each group of drill plans, the embodiment of the present application adds M groups of drill plans to the cache queue. By consuming the cache queue, each group of drill plans is exercised in turn to avoid concurrent execution of multiple groups of drill plans, thereby facilitating observation of the performance of the business system under each group of drill plans and improving test accuracy.

[0058] As an optional implementation, when performing a fault drill on the business system based on the execution strategy of the currently acquired drill plan and the case groups included in each group of drill plans, the above S164 may include the following steps: if the execution strategy of the currently acquired drill plan is sequential execution, selecting a case group from the currently acquired drill plan in sequence according to a first preset interval;

[0059] Based on the fault scenario cases in the currently selected case group, a drill task is generated and injected into the business system for fault drill.

[0060] The first preset interval time refers to a pre-set duration. After executing one case group, wait for a corresponding duration before executing the next case group. It should be understood that when executing each group of drill plans, the interval time can be set to be the same or different.

[0061] It should be understood that sequential execution refers to executing multiple case groups arranged serially within each drill plan based on the drill sequence during the planning phase. For example, a drill plan contains multiple case groups arranged serially, namely, Case Group A, Case Group B, and Case Group C. If the planned drill sequence is such that Case Group A executes after Case Group B, and Case Group C executes after Case Group A, then when executing this drill plan sequentially, Case Group B should be executed first, wait for the response time, then execute Case Group A, wait for the response time, and finally execute Case Group C.

[0062] The embodiment of the present application selects a case group from each group of drill plans in turn according to the drill sequence during arrangement, generates a drill task based on the fault case parameters contained in the fault scenario case in the case group, and injects it into the business system for fault drill, and waits for a corresponding period of time to better observe the performance of the business system under the case group, and then continues to select the next case group from the group of drill plans, generates a drill task and injects it into the business system for fault drill. In this way, when executing each drill plan, it is convenient to observe the performance of the system under each case group, improves the test accuracy, and at the same time improves the automation level of the fault drill. There is no need to manually inject fault scenario cases for drill, thereby improving execution efficiency and test accuracy.

[0063] As an optional implementation, when performing a fault drill on the business system based on the execution strategy of the currently acquired drill plan and the case groups included in each group of drill plans, the above S164 may include the following steps: if the execution strategy of the currently acquired drill plan is random execution, selecting a case group from the currently acquired drill plan based on a preset random algorithm at a second preset interval;

[0064] Based on the fault scenario cases in the currently selected case group, a drill task is generated and injected into the business system for fault drill.

[0065] The second preset interval time is similar to the first preset interval time mentioned above, and will not be repeated here.

[0066] Through a preset random algorithm, a case group is randomly selected from each group of drill plans, and based on the fault case parameters contained in the fault scenario cases in the case group, a drill task is generated and injected into the business system for fault drill, and a corresponding length of time is waited. Then, a case group is randomly selected from the group of drill plans, and a corresponding drill task is generated and injected into the business system for fault drill. In practical applications, this can be implemented using tools such as a random number generator and a random number seed, and the embodiments of the present application do not specifically limit this. For example: there is a group of drill plans containing multiple case groups arranged in a serial manner, namely case group A, case group B, and case group C. Assuming that it is a random selection without replacement, based on a preset random algorithm, the case group randomly selected from the group of drill plans for the first time is case group B and executed, and then a case group is randomly selected from the remaining case group B and case group C through a random algorithm and executed after the corresponding waiting time is met, until the case groups in the group of drill plans are executed.

[0067] When executing each group of drill plans, the embodiment of the present application randomly selects a case group from the currently obtained drill plan based on a preset random algorithm for drill to test the business system's response to sudden failures. For example, cases are selected from each group of drill plans with replacement through a random algorithm for drill. The number of selections can be preset, assuming 5. It is possible that the same case group may be selected 5 times in a row, and then the performance of the business system in the same fault situation for a long time can be observed, and then the performance of the business system in responding to various complex fault situations can be observed, thereby improving the accuracy of the test and enhancing the stability and security of the system.

[0068] As an optional implementation, the above S106 includes the following steps: determining an execution strategy for each group of drill plans based on the fault drill requirements, where the execution strategy includes at least one of the following execution modes: sequential execution and random execution;

[0069] Start M threads, each of which corresponds to M sets of exercise plans.

[0070] In each thread, based on the execution strategy of the corresponding drill plan and the case group included in the corresponding drill plan, a drill task is generated and injected into the business system for fault drill.

[0071] For example, one set of drill plans includes case group a, case group b, and case group c arranged in series and parallel, corresponding to random execution and thread 1; another set of drill plans includes case group b and case group c arranged in series, corresponding to sequential execution and thread 2. Start thread 1 and thread 2 to execute these two sets of drill plans respectively.

[0072] The embodiment of the present application starts multiple threads corresponding to each group of exercise plans and executes M groups of exercise plans at the same time. While improving the progress of the exercise, it can also ensure the orderly progress of each exercise plan. Even if one of the threads is disturbed, the remaining threads can proceed normally. It can also quickly discover potential problems in the business system and make timely adjustments to enhance the stability of the business system.

[0073] The following is a complete example to illustrate the fault drill method provided in the embodiment of the present application, which should not be understood as a limitation on the method in the embodiment of the present application.

[0074] First, if Figure 2 As shown, the fault drill is divided into four stages. In the first stage, fault scenario case 1, fault scenario case 2, fault scenario case 3 and fault scenario case 4 are selected from the fault scenario case set as target fault scenarios, that is, N=4; in the second stage, the target fault scenario cases to be drilled simultaneously are arranged in parallel in the same group to obtain concurrent orchestration group 1, concurrent orchestration group 2 and concurrent orchestration group 3 of the parallel orchestration layer, that is, T=3 case groups; in the third stage, the T case groups are serially orchestrated to obtain serial orchestration group 1 and serial orchestration group 2 of the serial orchestration layer, that is, M=2 groups of drill plans; in the fourth stage, the M groups of drill plans are executed to perform fault testing on the business system, including determining the execution strategy (i.e., sequential execution or random execution) and parameter settings (i.e., the first preset time or the second preset time) corresponding to each group of drill plans;

[0075] Among them, M groups of exercise plans can be executed serially. Figure 3A As shown, take out the most recent set of drill plans, assuming it is serial orchestration group 1, determine the execution strategy corresponding to the drill plan, assuming it is sequential execution, then take out the first experiment, that is, the first case group, that is, concurrent orchestration group 1, create and start the experiment, based on the fault case parameters of fault scenario case 1 and fault scenario case 2 contained in concurrent orchestration group 1, generate drill tasks and inject them into the business system for fault training at the same time, at equal intervals, that is, the first preset interval, continue to the next experiment, that is, concurrent orchestration group 2, based on the fault case parameters of fault scenario case 1, fault scenario case 2 and fault scenario case 3 contained in concurrent orchestration group 1, generate drill tasks and inject them into the business system for fault training at the same time, after serial orchestration group 1 is executed, continue to take out the next set of drill plans, that is, serial orchestration group 2;

[0076] It is also possible to execute M group exercise plans at the same time, such as Figure 3B As shown, all the drill plans are taken out and multiple threads are started. Each thread corresponds to one drill plan, that is, serial orchestration group 1 and serial orchestration group 2 correspond to two threads. These two threads are started at the same time, and the multiple concurrent orchestration groups contained therein are executed according to the execution strategy and interval time corresponding to each serial orchestration group.

[0077] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0078] With the above Figure 1 Corresponding to the fault drill method shown, the present application embodiment also provides a fault drill device. Figure 4 , is a structural diagram of a fault drill device 400 provided in an embodiment of the present application. The device 400 includes: an acquisition unit 410, an arrangement unit 420 and a drill unit 430.

[0079] The acquisition unit 410 is configured to acquire N target fault scenario cases from a fault scenario case set based on a fault drill requirement for the business system, and determine a drill order among the N target fault scenario cases;

[0080] The orchestration unit 420 is configured to orchestrate the N target fault scenario cases based on a preset orchestration strategy and the drill sequence to obtain M groups of drill plans, each group of drill plans including at least one fault scenario case, wherein the preset orchestration strategy includes at least one of the following orchestration modes: parallel orchestration and serial orchestration;

[0081] The drill unit 430 is configured to perform a fault drill on the business system based on the M groups of drill plans.

[0082] Optionally, when arranging the N target fault scenario cases based on the preset orchestration strategy and the drill sequence to obtain M groups of drill plans, the orchestration unit 420 performs the following steps: based on the drill sequence between the N target fault scenario cases, arranging the target fault scenario cases to be drilled simultaneously into the same group in parallel to obtain T case groups;

[0083] The T case groups are serially arranged to obtain the M group exercise plans.

[0084] Optionally, when each group of drill plans includes multiple case groups arranged in series, and each case group includes at least one target fault scenario case, the drill unit 430 performs the following steps when performing the fault drill on the business system based on the M groups of drill plans: determining an execution strategy for the currently obtained drill plan based on the fault drill requirements, the execution strategy including at least one of the following execution modes: sequential execution and random execution;

[0085] Add the M groups of drill plans to a cache queue;

[0086] The cache queue is consumed, and a fault drill is performed on the business system based on the execution strategy of the currently acquired drill plan and the case groups included in each group of drill plans.

[0087] Optionally, when performing a fault drill on the business system based on the execution strategy of the currently acquired drill plan and the case groups included in each group of drill plans, the drill unit 430 performs the following steps: if the execution strategy of the currently acquired drill plan is sequential execution, selecting a case group from the currently acquired drill plan in sequence according to a first preset interval time;

[0088] Based on the fault scenario cases in the currently selected case group, a drill task is generated and injected into the business system for fault drill.

[0089] Optionally, when performing a fault drill on the business system based on the execution strategy of the currently acquired drill plan and the case groups included in each group of drill plans, the drill unit 430 performs the following steps: if the execution strategy of the currently acquired drill plan is random execution, selecting a case group from the currently acquired drill plan based on a preset random algorithm according to a second preset interval;

[0090] Based on the fault scenario cases in the currently selected case group, a drill task is generated and injected into the business system for fault drill.

[0091] Optionally, when performing the fault drill on the business system based on the M groups of drill plans, the drill unit 430 performs the following steps: determining an execution strategy for each group of drill plans based on the fault drill requirements, wherein the execution strategy includes at least one of the following execution modes: sequential execution and random execution;

[0092] Start M threads, where the M threads correspond one-to-one to the M groups of exercise plans;

[0093] In each thread, based on the execution strategy of the corresponding drill plan and the case group included in the corresponding drill plan, a drill task is generated and injected into the business system to perform a fault drill.

[0094] Optionally, in the case where each target fault scenario case corresponds to a fault scenario, the fault drill apparatus 400 further includes a determination unit;

[0095] The determining unit is configured to, before the acquiring unit 410 acquires N target fault scenario cases from the fault scenario case set based on the fault drill requirement for the business system and determines the drill order between the N target fault scenario cases, perform the following steps: determining N fault scenarios corresponding to the business system based on the fault drill requirement;

[0096] For each fault scenario, split the fault scenario into multiple atomic scenarios according to the scenario category;

[0097] Obtain the fault case parameters corresponding to each atomic scenario from the atomic fault case set;

[0098] The fault case parameters of multiple atomic scenarios contained in the fault scenario are combined to obtain a target fault scenario case corresponding to the fault scenario and add it to the fault scenario case set.

[0099] Obviously, the fault drill device provided in the embodiment of the present application can be used as Figure 1 The execution subject of the fault drill method shown, for example Figure 1 In the fault drill method shown, step S102 can be performed by Figure 4 The acquisition unit 410 in the fault drill device shown in FIG. 1 is executed, and step S104 can be performed by Figure 4 The fault drill apparatus shown in FIG. 4 is executed by the arrangement unit 420, and step S106 can be performed by Figure 4 The fault rehearsal unit 430 in the fault rehearsal device is shown to execute.

[0100] According to another embodiment of the present application, Figure 4 The various units in the fault drill device shown can be individually or all combined into one or several other units to constitute, or one (or some) of the units can be further divided into multiple smaller functional units to constitute, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above-mentioned units are divided based on logical functions. In actual applications, the functions of one unit can also be implemented by multiple units, or the functions of multiple units can be implemented by one unit. In other embodiments of the present application, the fault drill device may also include other units. In actual applications, these functions can also be implemented with the assistance of other units, and can be implemented by the collaboration of multiple units.

[0101] According to another embodiment of the present application, the program can be executed on a general computing device such as a computer including a central processing unit (CPU), a random access memory (RAM), a read-only memory (ROM), and other processing elements and storage elements. Figure 1 A computer program (including program code) for each step of the corresponding method shown in FIG. Figure 4 The fault drill device shown in and the fault drill method of the embodiment of the present application are implemented. The computer program can be recorded on a computer-readable storage medium, for example, and transferred to an electronic device through the computer-readable storage medium and run therein.

[0102] Figure 5 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present application. Figure 5 At the hardware level, the electronic device includes a processor and, optionally, an internal bus, a network interface, and memory. The memory may include internal memory, such as high-speed random-access memory (RAM), or non-volatile memory, such as at least one disk drive. Of course, the electronic device may also include other hardware required for its services.

[0103] The processor, network interface, and memory can be interconnected via an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0104] The memory is used to store programs. Specifically, the program may include program code, which includes computer operating instructions. The memory may include internal memory and non-volatile memory, and provides instructions and data to the processor.

[0105] The processor reads the corresponding computer program from the non-volatile memory into the internal memory and then runs it, forming a fault drill device at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations:

[0106] Based on the fault drill requirements for the business system, N target fault scenario cases are obtained from the fault scenario case set, and a drill order between the N target fault scenario cases is determined;

[0107] Based on a preset orchestration strategy and the drill sequence, the N target fault scenario cases are orchestrated to obtain M groups of drill plans, each group of drill plans including at least one fault scenario case, the preset orchestration strategy including at least one of the following orchestration modes: parallel orchestration and serial orchestration;

[0108] Based on the M group drill plan, a fault drill is performed on the business system.

[0109] The above application Figure 4 The methods performed by the fault rehearsal device disclosed in the illustrated embodiments can be applied to a processor or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be performed by hardware integrated logic circuits in the processor or by software instructions. The above processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.

[0110] The electronic device may also perform Figure 1 Method, and realize the fault drill device in Figure 1 、 Figure 4 The functions of the illustrated embodiment will not be described in detail in the embodiments of the present application.

[0111] Of course, in addition to software implementation, the electronic device of this application does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0112] The embodiment of the present application also provides a computer-readable storage medium, which stores one or more programs, wherein the one or more programs include instructions, which, when executed by a portable electronic device including multiple application programs, can enable the portable electronic device to execute Figure 1 The method of the embodiment shown is specifically used to perform the following operations:

[0113] Based on the fault drill requirements for the business system, N target fault scenario cases are obtained from the fault scenario case set, and a drill order between the N target fault scenario cases is determined;

[0114] Based on a preset orchestration strategy and the drill sequence, the N target fault scenario cases are orchestrated to obtain M groups of drill plans, each group of drill plans including at least one fault scenario case, the preset orchestration strategy including at least one of the following orchestration modes: parallel orchestration and serial orchestration;

[0115] Based on the M group drill plan, a fault drill is performed on the business system.

[0116] In short, the above description is only a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.

[0117] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0118] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0119] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0120] The various embodiments in this specification are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the description of the method embodiments.

Claims

1. A fault drill method, characterized in that: include: Based on the fault drill requirements for the business system, N target fault scenario cases are obtained from the fault scenario case set, and a drill order between the N target fault scenario cases is determined; Based on a preset orchestration strategy and the drill sequence, the N target fault scenario cases are orchestrated to obtain M groups of drill plans, each group of drill plans including at least one fault scenario case, the preset orchestration strategy including at least one of the following orchestration modes: parallel orchestration and serial orchestration; Based on the M group drill plan, a fault drill is performed on the business system.

2. The method according to claim 1, characterized in that The N target fault scenario cases are arranged based on the preset arrangement strategy and the exercise sequence to obtain M groups of exercise plans, including: Based on the drill sequence between the N target fault scenario cases, the target fault scenario cases that are drilled simultaneously are arranged in parallel into the same group to obtain T case groups; The T case groups are serially arranged to obtain the M group exercise plans.

3. The method according to claim 1, characterized in that Each drill plan contains multiple case groups arranged in series, and each case group includes at least one target failure scenario case; The performing of a fault drill on the business system based on the M group drill plan includes: Based on the fault drill requirement, determining an execution strategy for the currently acquired drill plan, wherein the execution strategy includes at least one of the following execution modes: sequential execution and random execution; Add the M groups of drill plans to a cache queue; The cache queue is consumed, and a fault drill is performed on the business system based on the execution strategy of the currently acquired drill plan and the case groups included in each group of drill plans.

4. The method according to claim 3, characterized in that The execution strategy of the currently acquired drill plan and the case groups included in each drill plan are used to perform a fault drill on the business system, including: If the execution strategy of the currently acquired drill plan is sequential execution, then selecting a case group from the currently acquired drill plan in sequence according to the first preset interval time; Based on the fault scenario cases in the currently selected case group, a drill task is generated and injected into the business system for fault drill.

5. The method according to claim 3, characterized in that The execution strategy of the currently acquired drill plan and the case groups included in each drill plan are used to perform a fault drill on the business system, including: If the execution strategy of the currently acquired drill plan is random execution, then according to the second preset interval, a case group is selected from the currently acquired drill plan based on a preset random algorithm; Based on the fault scenario cases in the currently selected case group, a drill task is generated and injected into the business system for fault drill.

6. The method according to claim 1, wherein The performing of a fault drill on the business system based on the M group drill plan includes: Determine, based on the fault drill requirements, an execution strategy for each group of drill plans, wherein the execution strategy includes at least one of the following execution modes: sequential execution and random execution; Start M threads, where the M threads correspond one-to-one to the M groups of exercise plans; In each thread, based on the execution strategy of the corresponding drill plan and the case group included in the corresponding drill plan, a drill task is generated and injected into the business system to perform a fault drill.

7. The method according to claim 1, characterized in that Each target failure scenario case corresponds to a failure scenario; Before obtaining N target fault scenario cases from a fault scenario case set based on the fault drill requirement for the business system and determining a drill order among the N target fault scenario cases, the method further includes: Based on the fault drill requirements, determine N fault scenarios corresponding to the business system; For each fault scenario, split the fault scenario into multiple atomic scenarios according to the scenario category; Obtain the fault case parameters corresponding to each atomic scenario from the atomic fault case set; The fault case parameters of multiple atomic scenarios contained in the fault scenario are combined to obtain a target fault scenario case corresponding to the fault scenario and add it to the fault scenario case set.

8. A fault drill device, characterized in that: include: An acquisition unit is configured to acquire N target fault scenario cases from a fault scenario case set based on a fault drill requirement for the business system, and determine a drill order among the N target fault scenario cases; An orchestration unit is configured to orchestrate the N target fault scenario cases based on a preset orchestration strategy and the drill sequence to obtain M groups of drill plans, each group of drill plans including at least one fault scenario case, wherein the preset orchestration strategy includes at least one of the following orchestration modes: parallel orchestration and serial orchestration; A drill unit is used to perform a fault drill on the business system based on the M groups of drill plans.

9. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that When the instructions in the storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the method according to any one of claims 1 to 7.