Rehearsal method and device, computer equipment and storage medium
Patent Information
- Application Number
- CN202311237810.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-22
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-09-22
AI Technical Summary
[0003]有鉴于此,本发明提供了一种演练方法、装置、计算机设备及存储介质,以解决云平台测试不够充分的问题
Smart Images

Figure CN117194266B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and more specifically to training methods, apparatus, computer equipment, and storage media. Background Technology
[0002] With the development of information technology, cloud computing has gradually become a hot topic in the industry and plays an important role in various industries. Ensuring the robustness of cloud platforms is particularly important. However, current cloud platform testing is not sufficient in terms of fault testing and test cases. There are still many problems in the daily operation of cloud platforms, especially in fault scenarios. Summary of the Invention
[0003] In view of this, the present invention provides a training method, apparatus, computer equipment and storage medium to solve the problem of insufficient testing of cloud platforms.
[0004] In a first aspect, the present invention provides a training method, comprising: Obtain the training mode of the first training scenario, where the first training scenario is any one of all training scenarios to be trained; When the exercise mode is confirmed to be chaotic mode, both the fault exercise thread and the test case exercise thread are created simultaneously. The pre-configured fault execution parameters are obtained from the fault drill thread, and fault instructions are generated based on the fault execution parameters. The test cases to be executed are obtained from the test case drill thread; The fault instructions and test cases are sent synchronously to the target cloud platform so that the target cloud platform can conduct fault drills based on the fault instructions and test case drills based on the test cases.
[0005] The above method is used to obtain the training mode of the first training scenario, which is any one of the training scenarios to be trained. When the training mode is confirmed to be chaotic mode, a fault training thread and a test case training thread are created simultaneously. The fault training thread obtains pre-configured fault execution parameters and generates fault instructions based on these parameters. The test case training thread obtains test cases to be executed. The fault instructions and test cases are synchronously sent to the target cloud platform so that the target cloud platform can perform fault training based on the fault instructions and test case training simultaneously. In chaotic mode, fault training and test case training can be performed simultaneously, fully combining the advantages of both test case testing and fault testing, and offering greater advantages than either alone. Compared to ordinary test case testing, chaotic mode adds fault injection, which is more conducive to verifying the robustness of functional test cases and the robustness of the cloud platform. Compared to ordinary fault testing, chaotic mode adds test case execution, which is beneficial for verifying the impact of specific faults on specific test cases, increasing the flexibility and accuracy of testing. In addition, both fault training and test case training in this method are automated, reducing labor costs and improving execution efficiency. It can inject faults more accurately and create chaos, and find vulnerabilities in cloud platforms that are not easy to find in a targeted manner; and reproduce, analyze and deal with faults that are difficult to check and verify manually in an automated way, so as to discover vulnerabilities in cloud platforms more accurately and efficiently and improve the robustness of cloud platforms.
[0006] In one optional implementation, both the fault simulation thread and the test case simulation thread include the same pairing identifier, and the method further includes: Monitor the status of the first target thread, which includes either the fault simulation thread or the test case simulation thread; When the status of the first target thread is detected as stopped, the second target thread is searched for in the threads created in the chaos mode according to the pairing identifier, and the exercise operation of the second target thread is stopped. The second target thread is the exercise thread created in the chaos mode other than the first target thread.
[0007] In this way, fault drill threads and test case drill threads belonging to the same drill scenario can be associated through pairing identifiers, ensuring that the drills can be stopped synchronously.
[0008] In one optional implementation, when the drill mode is confirmed to be chaotic mode, after creating both the fault drill thread and the test case drill thread, the method further includes: Store the pairing identifier, the first thread information of the fault simulation thread, and the second thread information of the test case simulation thread in the same global variable.
[0009] In one optional implementation, when the state of the first target thread is detected as stopped, the second target thread is located among the threads created in the chaotic mode according to the pairing identifier, and the exercise operation of the second target thread is stopped, including: The thread information of the second target thread is retrieved from the global variables based on the pairing identifier; Determine the second target thread based on the thread information; Send a stop signal to the second target thread; The execution status of the second target thread is updated to the stopped state based on the stop signal; When the execution status of the second target thread is detected as stopped, the second target thread is controlled to enter the stop routine. The stop routine is used to control the second target to perform a stop drill operation. Wherein, when the first target thread includes a test case exercise thread, the second target thread includes a fault exercise thread corresponding to the test case exercise thread; or, When the first target thread includes a fault simulation thread, the second target thread includes a test case simulation thread corresponding to the fault simulation thread.
[0010] Using the above method, a stop signal can be sent to another thread to control the corresponding thread to stop the program and exit the exercise, thus ensuring the synchronization of the two threads.
[0011] In one alternative implementation, the thread information is the address of a thread object reference.
[0012] In one optional implementation, the fault execution parameters include fault scenarios and fault operation parameters corresponding to each fault scenario; the pre-configured fault execution parameters are obtained according to the fault drill thread, and fault instructions are generated based on the fault execution parameters, including: The pre-configured fault scenarios and the corresponding fault scenario parameters for each fault scenario are obtained from the fault drill thread. A first fault instruction is generated based on the first fault scenario and the fault scenario parameters corresponding to the first fault scenario, wherein the first fault scenario is any fault scenario.
[0013] In one optional implementation, the drill mode includes one or more of the following: fault drill mode, use case drill mode, and chaos drill mode, wherein... When the exercise mode is confirmed to be test case mode, obtain the pre-configured test cases; Send the test cases to the target cloud platform so that the target cloud platform can perform test case drills based on the test cases; or, When the exercise mode is confirmed to be fault mode, the pre-configured fault execution parameters are obtained, and fault instructions are generated based on the fault execution parameters. The fault command is sent to the target cloud platform so that the target cloud platform can conduct fault drills based on the fault command.
[0014] Using the methods described above, test case drills or fault drills can also be performed separately, making the drill methods more flexible.
[0015] Secondly, the present invention provides a training apparatus, comprising: The first acquisition module is used to acquire the training mode of the first training scenario, wherein the first training scenario is any one of all training scenarios to be trained; Create a module to simultaneously create a fault drill thread and a test case drill thread when the drill mode is confirmed to be chaotic mode; The processing module is used to obtain pre-configured fault execution parameters from the fault drill thread and generate fault instructions based on the fault execution parameters; The second acquisition module is used to acquire test cases to be executed based on the test case exercise thread; The first sending module is used to synchronously send fault instructions and test cases to the target cloud platform, so that the target cloud platform can perform fault drills according to the fault instructions and test case drills at the same time.
[0016] In some alternative embodiments, the apparatus further includes: The monitoring module is used to monitor the status of the first target thread, which includes either the fault simulation thread or the test case simulation thread. The lookup module is used to find the second target thread in the threads created in the chaos mode according to the pairing identifier when the status of the first target thread is stopped, and to stop the exercise operation of the second target thread. The second target thread is the exercise thread created in the chaos mode other than the first target thread.
[0017] In some alternative embodiments, the apparatus further includes: The storage module is used to store the pairing identifier, the first thread information of the fault simulation thread, and the second thread information of the test case simulation thread into the same global variable.
[0018] In some optional implementations, when the status of the first target thread is detected as stopped during the exercise, the lookup module includes: The lookup unit is used to search for the thread information of the second target thread from global variables based on the pairing identifier; The determining unit is used to determine the second target thread based on thread information. The sending unit is used to send a stop signal to the second target thread; The update unit is used to update the execution status of the second target thread to a stopped state based on the stop signal; The stopping unit is used to control the second target thread to enter the stopping procedure when the execution status of the second target thread is detected to be stopped. The stopping procedure is used to control the second target to perform a stop exercise operation. When the first target thread includes a test case exercise thread, the second target thread includes a fault exercise thread corresponding to the test case exercise thread; or, when the first target thread includes a fault exercise thread, the second target thread includes a test case exercise thread corresponding to the fault exercise thread.
[0019] In some alternative implementations, the thread information in the lookup unit is the thread object reference address.
[0020] In some optional implementations, the fault execution parameters include fault scenarios and fault operation parameters corresponding to each fault scenario; the processing module includes: The acquisition unit is used to acquire pre-configured fault scenarios and fault scenario parameters corresponding to each fault scenario based on the fault drill thread. The generation unit is used to generate a first fault instruction based on a first fault scenario and the fault scenario parameters corresponding to the first fault scenario, wherein the first fault scenario is any fault scenario.
[0021] In some optional implementations, the drill modes include one or more of fault drill modes, use case drill modes, and chaos drill modes, and the apparatus further includes: The third acquisition module is used to acquire pre-configured test cases when the exercise mode is confirmed to be test case mode; The second sending module is used to send test cases to the target cloud platform so that the target cloud platform can perform test case drills based on the test cases; or, The fourth acquisition module is used to acquire pre-configured fault execution parameters and generate fault instructions based on the fault execution parameters when the exercise mode is confirmed to be fault mode. The third sending module is used to send fault commands to the target cloud platform so that the target cloud platform can conduct fault drills based on the fault commands.
[0022] Thirdly, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to perform the exercise method described in the first aspect or any corresponding embodiment thereof.
[0023] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to perform the exercise method of the first aspect or any corresponding embodiment thereof. Attached Figure Description
[0024] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0025] Figure 1 This is a flowchart illustrating the exercise method according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating a method for stopping a second target thread according to an embodiment of the present invention; Figure 3 This is a flowchart illustrating another method for stopping a second target thread according to an embodiment of the present invention; Figure 4 This is a flowchart illustrating another training method according to an embodiment of the present invention; Figure 5 This is a schematic diagram of another training method according to an embodiment of the present invention; Figure 6 This is a flowchart illustrating another training method according to an embodiment of the present invention; Figure 7 This is a structural block diagram of the training device according to an embodiment of the present invention; Figure 8 This is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0027] With the development of information technology, cloud computing has gradually become a hot topic in the industry, and major domestic and foreign vendors have begun to deploy their cloud computing service platforms in various fields. Cloud computing includes the following service layers: Infrastructure as a Service (IaaS), Platform as a Service (PaaS), and Software as a Service (SaaS). The cloud platform, serving as the management platform for cloud computing services at the IaaS layer, plays a crucial role. The cloud platform manages a large number of virtual machines, supporting various functionalities and services. It also needs to handle numerous failure scenarios, including hard drive failures, excessive storage pool I / O pressure, virtual machine operating system crashes, network card failures, switch failures, network lag, server power outages, and combinations thereof. Therefore, ensuring the robustness of the cloud platform system is particularly important.
[0028] In a complex system, human intervention alone cannot prevent all failures. Instead, we should focus on identifying as many vulnerable and fault-prone components as possible that could lead to these anomalies before they are triggered. Once these risks are identified, the system can be hardened and protected in a targeted manner to avoid the serious consequences of failures. However, current testing in this area is insufficient.
[0029] Based on this, according to an embodiment of the present invention, a training method embodiment is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0030] This embodiment provides a training method that can be used with the aforementioned computer equipment, such as servers and cloud platforms. Figure 1 This is a flowchart of a training method according to an embodiment of the present invention, such as... Figure 1 As shown, the process includes the following steps: Step S101: Obtain the training mode of the first training scenario.
[0031] Specifically, the first training scenario is any one of the scenarios to be trained.
[0032] In one optional example, cloud platform drills can be conducted simultaneously. Each drill scenario can select a drill mode, such as test case drill mode, fault drill mode, or chaos mode. Test case drill mode involves executing prescribed test cases in a normal cloud platform environment. Fault drill mode involves selecting one or more fault scenarios, such as host CPU load or host memory load, and injecting the selected faults into the cloud platform. Chaos mode involves executing fault drills and test case drills simultaneously within the same time period.
[0033] In an optional example, chaos engineering can be used, which allows for the configuration of fault simulation mode, test case simulation mode, and chaos mode. Users can simply select the desired mode for simulation.
[0034] Step S102: When it is confirmed that the exercise mode is chaotic mode, a fault exercise thread and a test case exercise thread are created at the same time.
[0035] Specifically, when the exercise mode is confirmed to be chaotic mode, two independent threads can be created simultaneously through an asynchronous executor: a fault exercise thread and a test case exercise thread. The fault exercise thread is used to execute fault exercises, and the test case exercise thread is used to execute test case exercises, ensuring the synchronization of their start-up.
[0036] Step S103: Obtain the pre-configured fault execution parameters according to the fault drill thread, and generate fault instructions based on the fault execution parameters.
[0037] Specifically, the fault simulation thread obtains pre-configured fault execution parameters, such as fault scenario parameters (host CPU load, host power failure, etc.) and fault operation parameters (e.g., memory load percentage). Fault instructions are then generated based on these parameters to instruct the cloud platform to perform fault simulations.
[0038] In an optional example, such as using a chaos engineering control cloud platform for drills, selecting the fault drill mode, the steps could be as follows: Step a1: Select the target cloud platform for the exercise; there can be one or more. Step a2: Select fault scenario parameters, such as host CPU load, host memory load, host power failure, etc. Step a3: Configure fault operation parameters, such as the percentage of CPU load and the percentage of memory load.
[0039] Fault instructions are generated based on the target cloud platform, fault scenario parameters, and fault operation parameters.
[0040] Step S104: Obtain the test cases to be executed according to the test case drill thread.
[0041] Specifically, the test case drill thread is used to obtain the test cases to be executed.
[0042] In an optional example, for instance, the target cloud platform is determined, which is the same as the cloud platform used to perform the fault simulation. The set of test cases for the preparation phase is obtained. When it is necessary to determine the number of executions, the number of executions in the preparation phase can be obtained, or it can be set to execute once if the number of executions is not obtained. The set of test cases for the running phase is obtained, as well as the set of test cases for the cleanup phase are obtained.
[0043] Step S105: Simultaneously send the fault command and test cases to the target cloud platform so that the target cloud platform can conduct fault drills based on the fault command and test case drills based on the test cases.
[0044] Specifically, fault instructions and test cases are sent synchronously to the target cloud platform. The cloud platform then conducts fault drills based on the fault instructions and test case drills based on the test cases.
[0045] In an optional example, a test case execution command can also be sent, and the target cloud platform can obtain the required test cases based on the test case execution command to execute the test case exercise.
[0046] In an alternative example, the command module of chaos engineering can be used to send fault commands to the probe module of the target cloud platform. The probe module then converts the fault commands into local commands to execute fault drills. The command module obtains the progress and results of the fault drills by periodically querying the probes and receiving callbacks.
[0047] The exercise method provided in this embodiment obtains the exercise mode of a first exercise scenario, wherein the first exercise scenario is any one of all exercise scenarios to be exercised; when the exercise mode is confirmed to be chaotic mode, a fault exercise thread and a test case exercise thread are created simultaneously; pre-configured fault execution parameters are obtained according to the fault exercise thread, and fault instructions are generated based on the fault execution parameters; test cases to be executed are obtained according to the test case exercise thread; the fault instructions and test cases are synchronously sent to the target cloud platform so that the target cloud platform can perform fault exercise according to the fault instructions and test case exercise according to the test cases. In chaotic mode, fault exercise and test case exercise can be performed simultaneously, fully combining the advantages of both test case testing and fault testing, and is more advantageous than either one alone. Compared to standard test cases, chaos mode adds fault injection, which is more conducive to verifying the robustness of functional test cases and the robustness of the cloud platform. Compared to standard fault testing, chaos mode adds test case execution, which is beneficial for verifying the impact of specific faults on specific test cases, increasing the flexibility and accuracy of testing. In addition, both fault simulation and test case simulation in this method are executed automatically, which can reduce labor costs and improve execution efficiency. It can more accurately inject faults and create chaos, and specifically look for vulnerabilities in the cloud platform that are not easy to find. Furthermore, it can reproduce, analyze, and respond to faults that are difficult to verify manually in an automated way, which can more accurately and efficiently discover vulnerabilities in the cloud platform and improve the robustness of the cloud platform.
[0048] In one optional implementation, both the fault simulation thread and the test case simulation thread include the same pairing identifier, and the method further includes, for example, Figure 2 The steps shown are as follows: Step S201: Monitor the status of the first target thread.
[0049] Specifically, the first target thread includes either a fault simulation thread or a test case simulation thread. A monitoring program can be used to monitor the status of these threads to ensure they are functioning correctly.
[0050] Step S202: When the status of the first target thread is detected as stopped, the second target thread is searched in the threads created in the chaos mode according to the pairing identifier, and the exercise operation of the second target thread is stopped.
[0051] Specifically, the second target thread is the exercise thread created in the chaos mode, excluding the first target thread. When a test case exercise thread stops, the corresponding fault exercise thread is located from the created threads based on the pairing identifier, and the fault exercise is stopped. Alternatively, when a fault exercise thread is detected, the corresponding test case exercise thread is located from the created threads based on the pairing identifier, and the test case exercise is stopped.
[0052] In an optional example, the pairing identifier can be the parent ID of the thread. Test case exercise threads belonging to the same chaotic mode in the same exercise scenario are associated with the same parent ID, and the corresponding thread can be found by querying the parent ID.
[0053] In an optional implementation, after confirming that the drill mode is chaotic mode and creating both the fault drill thread and the test case drill thread, the method further includes: Store the pairing identifier, the first thread information of the fault simulation thread, and the second thread information of the test case simulation thread in the same global variable.
[0054] Specifically, once the exercise mode is confirmed to be chaotic mode, and both the fault simulation thread and the test case simulation thread are created, the pairing identifier, the first thread information of the fault simulation thread, and the second thread information of the test case simulation thread are stored in the same global variable. This global variable serves as the "bridge" for communication between the fault simulation thread and the test case simulation thread. The thread information of the required thread can then be queried through this global variable.
[0055] In an optional example, when multiple drill scenarios are in chaotic mode, the global variable will contain multiple pairs of fault drill and test case drill thread information, with each pair of thread information associated with the same parent ID.
[0056] In one optional implementation, when the state of the first target thread is detected as stopped, the second target thread is located among the threads created in chaotic mode based on the pairing identifier, and the exercise operation of the second target thread is stopped, including as follows: Figure 3 The steps shown are as follows: Step S301: Search for the thread information of the second target thread from the global variables based on the pairing identifier.
[0057] Specifically, the pairing identifier can be the parent ID, or other identifiers such as a specific identifier name, as long as it can uniquely identify the corresponding second target thread. The thread information of the second target thread is then retrieved from global variables based on the parent ID.
[0058] In an optional example, in chaotic mode, test case execution can be configured to not stop testing even if the test case fails. Testing will only stop when the test is manually stopped, when the test case completes execution, or when a stop signal is received from the fault simulation thread. This is because in chaotic mode, the probability of test case failure is greatly increased. If testing stops when the test case fails, most test cases will not be executed, and testing will not be sufficient.
[0059] In one alternative implementation, the thread information is the address of a thread object reference.
[0060] Specifically, the thread object reference address is used to store the thread object, and the target thread can be located through the thread object reference address.
[0061] Step S302: Determine the second target thread based on the thread information.
[0062] Specifically, the second target thread is located based on the thread object reference address.
[0063] Step S303: Send a stop signal to the second target thread.
[0064] Specifically, a stop signal is sent to the second target thread. Because the second target thread cannot be interrupted immediately and requires a stop procedure to execute the stop exercise, an immediate interruption could alter the cloud platform environment, necessitating manual restoration to the original environment and causing inconvenience. Therefore, a stop signal can be sent to the second target thread.
[0065] In an optional example, the stop signal can be an interrupt signal.
[0066] Step S304: Update the execution status of the second target thread to the stopped state based on the stop signal.
[0067] Step S305: When the execution status of the second target thread is detected to be stopped, the second target thread is controlled to enter the stop procedure. The stop procedure is used to control the second target to perform a stop exercise operation.
[0068] Specifically, when the first target thread includes a test case simulation thread, the second target thread includes a fault simulation thread corresponding to the test case simulation thread; or, when the first target thread includes a fault simulation thread, the second target thread includes a test case simulation thread corresponding to the fault simulation thread. The execution status of the second target thread is updated to stopped based on the stop signal. When the execution status of the second target thread is detected to be stopped, the second target thread is controlled to stop the corresponding fault simulation or test case simulation.
[0069] In one optional implementation, the fault execution parameters include fault scenarios and fault operation parameters corresponding to each fault scenario; the pre-configured fault execution parameters are obtained according to the fault drill thread, and fault instructions are generated based on the fault execution parameters, including: Step b1: Obtain the pre-configured fault scenarios and the fault scenario parameters corresponding to each fault scenario based on the fault drill thread.
[0070] Specifically, fault scenarios can be obtained in the form of fault scenario parameters, such as host CPU load, host memory load, host power failure, etc. Fault operation parameters include runtime cloud platform environment data, such as the percentage of CPU load reaching the CPU fault threshold, such as 80%, and the percentage of memory load reaching the memory fault threshold.
[0071] Step b2: Generate a first fault instruction based on the first fault scenario and the fault scenario parameters corresponding to the first fault scenario.
[0072] Specifically, the first fault scenario is any fault scenario, and the fault instruction contains the fault scenario and fault operation parameters.
[0073] In one optional implementation, the drill mode includes one or more of the following: fault drill mode, use case drill mode, and chaos drill mode, wherein... Step c1: When the exercise mode is confirmed to be test case mode, obtain the pre-configured test cases.
[0074] Step c2: Send the test cases to the target cloud platform so that the target cloud platform can perform test case drills based on the test cases.
[0075] Specifically, the exercise mode can be configured as test case mode. When the exercise mode is confirmed to be test case mode, the test cases to be executed are obtained, including the test case set for the preparation phase, the test case set for the configuration and operation phase, and the test case set for the configuration and cleanup phase. The number of test case executions can also be obtained, and all test cases are sent to the target platform.
[0076] In an optional example, the execution order and execution scenario of test cases can also be set. By obtaining the execution time nodes, the test cases to be executed can be determined. For example, if the execution time node is the preparation phase, the test cases in the preparation phase can be obtained and executed until the test cases at all time nodes have been executed.
[0077] or, Step c3: When the exercise mode is confirmed to be fault mode, obtain the pre-configured fault execution parameters and generate fault instructions based on the fault execution parameters.
[0078] Step c4: Send the fault command to the target cloud platform so that the target cloud platform can conduct fault drills based on the fault command.
[0079] Specifically, when configured in fault mode, only fault drills are performed to obtain pre-configured fault execution parameters, generate fault instructions based on the fault execution parameters, and send the fault instructions to the target cloud platform. The target cloud platform then performs fault drills based on the fault execution parameters upon receiving the fault instructions.
[0080] To make the method of this invention clearer, this invention also provides a specific embodiment of the exercise method, for example, encapsulating the various functions of chaos engineering into functional modules, including an exercise definition module, a fault exercise execution module, a test case exercise execution module, and an exercise result feedback module, such as... Figure 4 As shown, the exercise definition module is used to define the exercise execution strategy (including fault exercise mode, test case exercise mode, or chaos mode), configure exercise parameters, and create the exercise. The configuration items will vary depending on the exercise execution strategy.
[0081] If the exercise execution strategy is fault drill mode, define the fault drill as follows: 1. Select the target cloud platform for the exercise, which can be one or more; 2. Select the fault scenario, such as host CPU load, host memory load, host power failure, etc.; 3. Configure fault parameters, such as the percentage of CPU load and the percentage of memory load.
[0082] If the exercise execution strategy is the use case exercise mode, define the use case exercise as follows: 1. Select the target cloud platform for the exercise, which can be one or more; 2. Configure the use case set for the preparation phase and the number of times the preparation phase is executed; 3. Configure the use case set for the runtime phase; 4. Configure the use case set for the cleanup phase; 5. Configure the use case parameters, such as specifying the operation object of the use case as a specific virtual machine.
[0083] If the exercise execution strategy is chaotic mode, define fault drills and test case drills: 1. Select the target cloud platform for the exercise, which can be one or more; 2. Select the fault scenario, such as host CPU load, host memory load, host power failure, etc., for the execution module to inject the corresponding fault on the target host; 3. Configure fault parameters, such as the percentage of CPU load and the percentage of memory load; 4. Configure the test case set for the preparation phase and the number of executions in the preparation phase; 5. Configure the test case set for the runtime phase; 6. Configure the test case set for the cleanup phase; 7. Configure test case parameters, such as specifying the operation object of the test case as a specific virtual machine.
[0084] The fault drill execution module is used to execute fault drills defined in the drill definition module. The prerequisite for entering this module is that the drill execution strategy is either the default mode or the chaotic mode. The fault drill execution module consists of an instruction module located in the chaotic engineering system and probe modules pre-installed on all hosts of the target cloud platform, as shown in the diagram. Figure 5As shown, the instruction module sends fault instructions to the probe module, which then converts these instructions into local commands and executes them on the host. The instruction module periodically queries and receives callbacks from the probes to obtain the progress and results of the fault drill execution. Through this coordination, the fault drill execution module controls the injection and recovery of faults configured in the drill definition module on the cloud platform. If a stop is triggered due to an anomaly or manual intervention, and there are related use case drills, an interrupt signal (stop signal) will be sent to their corresponding threads simultaneously.
[0085] The test case execution module is used to execute test cases defined in the test case definition module. The prerequisite for entering this module is that the execution strategy is either a long-term stable mode or a chaotic mode. A test case can be viewed as a collection of one or more interfaces of the cloud platform. This module uses the cloud platform's SDK to call the relevant interfaces and execute the test cases in the aforementioned defined set. Furthermore, it retrieves the progress and results of the test case execution through the cloud platform's query interface. If an execution is stopped due to an exception or manually triggered, and there are related fault test cases, an interrupt signal (stop signal) will be sent to their corresponding threads simultaneously.
[0086] The exercise result feedback module is used to persistently record the exercise execution results returned by the fault exercise execution module and the test case exercise execution module to the database, and push the feedback to the front-end interface of the chaos engineering system.
[0087] The specific process is as follows Figure 6 As shown, select the exercise strategy and judge the exercise strategy. If it is the fault exercise mode, execute the fault exercise. If it is the test case exercise mode, execute the test case exercise. If it is neither the fault exercise mode nor the test case exercise mode, it is confirmed as the chaos mode. Simultaneously execute the fault exercise and the test case exercise, record the exercise execution results and feed them back to the front-end interface of the chaos engineering system.
[0088] This embodiment also provides a training device for implementing the above embodiments and preferred embodiments, and details already described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0089] This embodiment provides a training device, such as Figure 7 As shown, it includes: The first acquisition module 701 is used to acquire the training mode of the first training scenario, wherein the first training scenario is any one of all training scenarios to be trained; Create module 702 to simultaneously create a fault drill thread and a test case drill thread when the drill mode is confirmed to be chaotic mode; The processing module 703 is used to obtain pre-configured fault execution parameters according to the fault drill thread, and generate fault instructions based on the fault execution parameters; The second acquisition module 704 is used to acquire test cases to be executed based on the test case exercise thread; The first sending module 705 is used to synchronously send fault instructions and test cases to the target cloud platform, so that the target cloud platform can perform fault drills according to the fault instructions and test case drills at the same time.
[0090] In some alternative embodiments, the apparatus further includes: The monitoring module 706 is used to monitor the status of the first target thread, which includes a fault drill thread or a test case drill thread. The lookup module 707 is used to find the second target thread in the threads created in the chaos mode according to the pairing identifier when the status of the first target thread is stopped, and to stop the exercise operation of the second target thread. The second target thread is an exercise thread created in the chaos mode other than the first target thread.
[0091] In some alternative embodiments, the apparatus further includes: Storage module 708 is used to store the pairing identifier, the first thread information of the fault simulation thread, and the second thread information of the test case simulation thread into the same global variable.
[0092] In some optional implementations, when the status of the first target thread is detected as stopped, the lookup module 707 includes: The lookup unit is used to search for the thread information of the second target thread from global variables based on the pairing identifier; The determining unit is used to determine the second target thread based on thread information. The sending unit is used to send a stop signal to the second target thread; The update unit is used to update the execution status of the second target thread to a stopped state based on the stop signal; The stopping unit is used to control the second target thread to enter the stopping procedure when the execution status of the second target thread is detected to be stopped. The stopping procedure is used to control the second target to perform a stop exercise operation. When the first target thread includes a test case exercise thread, the second target thread includes a fault exercise thread corresponding to the test case exercise thread; or, when the first target thread includes a fault exercise thread, the second target thread includes a test case exercise thread corresponding to the fault exercise thread.
[0093] In some alternative implementations, the thread information in the lookup unit is the thread object reference address.
[0094] In some optional implementations, the fault execution parameters include fault scenarios and fault operation parameters corresponding to each fault scenario; the processing module 703 includes: The acquisition unit is used to acquire pre-configured fault scenarios and fault scenario parameters corresponding to each fault scenario based on the fault drill thread. The generation unit is used to generate a first fault instruction based on a first fault scenario and the fault scenario parameters corresponding to the first fault scenario, wherein the first fault scenario is any fault scenario.
[0095] In some optional implementations, the drill modes include one or more of fault drill modes, use case drill modes, and chaos drill modes, and the apparatus further includes: The third acquisition module 709 is used to acquire pre-configured test cases when the exercise mode is confirmed to be test case mode; The second sending module 710 is used to send test cases to the target cloud platform so that the target cloud platform can perform test case drills based on the test cases. or, The fourth acquisition module 711 is used to acquire pre-configured fault execution parameters and generate fault instructions based on the fault execution parameters when the exercise mode is confirmed to be fault mode. The third sending module 712 is used to send fault instructions to the target cloud platform so that the target cloud platform can perform fault drills based on the fault instructions.
[0096] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.
[0097] In this embodiment, the training device is presented in the form of a functional unit. Here, a unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0098] This invention also provides a computer device having the above-described features. Figure 7 The training apparatus shown.
[0099] Please see Figure 8 , Figure 8 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of the present invention, such as... Figure 8As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 8 Take a processor 10 as an example.
[0100] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.
[0101] The memory 20 stores instructions executable by at least one processor 10 to cause at least one processor 10 to perform the method shown in the above embodiments.
[0102] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0103] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.
[0104] The computer device also includes an input device 30 and an output device 40. The processor 10, memory 20, input device 30, and output device 40 can be connected via a bus or other means. Figure 8 Taking the example of a connection between China and Israel via a bus.
[0105] Input device 30 can receive input numerical or character information, and generate key signal inputs related to user settings and function control of the computer device, such as a touchscreen, keypad, mouse, trackpad, touchpad, joystick, one or more mouse buttons, trackball, joystick, etc. Output device 40 may include display devices, auxiliary lighting devices (e.g., LEDs), and haptic feedback devices (e.g., vibration motors). The aforementioned display devices include, but are not limited to, liquid crystal displays, light-emitting diodes, displays, and plasma displays. In some alternative embodiments, the display device may be a touchscreen.
[0106] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.
[0107] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A training method, characterized in that, The method includes: Obtain the training mode of the first training scenario, where the first training scenario is any one of all training scenarios to be trained; When the exercise mode is confirmed to be chaotic mode, a fault exercise thread and a test case exercise thread are created simultaneously. The fault drill thread obtains pre-configured fault execution parameters and generates fault instructions based on the fault execution parameters. The test cases to be executed are obtained from the test case drill thread. The fault command and the test cases are sent synchronously to the target cloud platform so that the target cloud platform can perform fault drills based on the fault command and test case drills based on the test cases.
2. The method according to claim 1, characterized in that, Both the fault simulation thread and the test case simulation thread include the same pairing identifier, and the method further includes: Monitor the status of the first target thread, wherein the first target thread includes the fault simulation thread or the test case simulation thread; When the monitoring shows that the state of the first target thread is stopped, the second target thread is searched in the threads created in the chaos mode according to the pairing identifier, and the exercise operation of the second target thread is stopped. The second target thread is an exercise line created in the chaos mode other than the first target thread.
3. The method according to claim 2, characterized in that, When the exercise mode is confirmed to be chaotic mode, after creating both the fault exercise thread and the test case exercise thread, the method further includes: The pairing identifier, the first thread information of the fault simulation thread, and the second thread information of the test case simulation thread are stored in the same global variable.
4. The method according to claim 3, characterized in that, When the status of the first target thread is detected as "stopped rehearsal," the second target thread is located in the threads created in the chaotic mode according to the pairing identifier, and the rehearsal operation of the second target thread is stopped, including: The thread information of the second target thread is retrieved from the global variables based on the pairing identifier; The second target thread is determined based on the thread information; Send a stop signal to the second target thread; The execution status of the second target thread is updated to a stopped state based on the stop signal; When the execution status of the second target thread is detected to be stopped, the second target thread is controlled to enter a stop procedure. The stop procedure is used to control the second target to stop the exercise operation. Wherein, when the first target thread includes the test case exercise thread, the second target thread includes the fault exercise thread corresponding to the test case exercise thread; or, When the first target thread includes the fault simulation thread, the second target thread includes the test case simulation thread corresponding to the fault simulation thread.
5. The method according to claim 4, characterized in that, The thread information is the reference address of the thread object.
6. The method according to any one of claims 2 to 5, characterized in that, The fault execution parameters include fault scenarios and fault operation parameters corresponding to each fault scenario; the step of obtaining pre-configured fault execution parameters according to the fault drill thread and generating fault instructions based on the fault execution parameters includes: The fault simulation thread obtains the pre-configured fault scenarios and the fault scenario parameters corresponding to each fault scenario. A first fault instruction is generated based on a first fault scenario and the fault scenario parameters corresponding to the first fault scenario, wherein the first fault scenario is any fault scenario.
7. The method according to any one of claims 2 to 5, characterized in that, The exercise modes include one or more of the following: fault simulation mode, use case simulation mode, and chaos simulation mode. When the exercise mode is confirmed to be test case mode, obtain the pre-configured test cases; The test cases are sent to the target cloud platform so that the target cloud platform can perform test case drills based on the test cases; or, When the exercise mode is confirmed to be a fault mode, the pre-configured fault execution parameters are obtained, and a fault instruction is generated based on the fault execution parameters. The fault command is sent to the target cloud platform so that the target cloud platform can perform fault drills based on the fault command.
8. A training device, characterized in that, The device includes: The first acquisition module is used to acquire the training mode of the first training scenario, wherein the first training scenario is any one of all training scenarios to be trained; A module is created to simultaneously create a fault drill thread and a test case drill thread when the drill mode is confirmed to be chaotic mode. The processing module is used to obtain pre-configured fault execution parameters according to the fault drill thread, and generate fault instructions based on the fault execution parameters; The second acquisition module is used to acquire test cases to be executed based on the test case exercise thread; The sending module is used to synchronously send the fault instruction and the test case to the target cloud platform, so that the target cloud platform can perform fault drills according to the fault instruction and test case drills according to the test case.
9. A computer device, characterized in that, include: A memory and a processor are communicatively connected, the memory storing computer instructions, and the processor executing the computer instructions to perform the exercise method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to perform the exercise method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Distributed system test method and system based on chaos experiment
CN110765023A
Drill method and system applied to chaos engineering
CN114647489A