Chaotic engineering fault drilling method and device, storage medium and electronic equipment
By dynamically adjusting the fault factor in chaos engineering, combining performance indicators and feedback from user experience dimensions, the low accuracy problem caused by static fault injection strategies is solved, and a more efficient fault drill is achieved.
Patent Information
- Application Number
- CN202510645988.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-08-15
AI Technical Summary
In traditional chaos engineering fault drills, static fault injection strategies lead to low accuracy of drill results, and it is impossible to provide real-time feedback and adjust experiments.
By injecting fault factors and obtaining response results, the fault factors are dynamically adjusted to adapt to the system state, and real-time feedback and adjustments are performed using performance metrics and user experience dimensions.
It improves the accuracy of chaos engineering fault drills, realizes real-time monitoring and dynamic adjustment of system status, and improves the accuracy and efficiency of experiments.
Smart Images

Figure CN120498958A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of chaos engineering, and in particular to a chaos engineering fault drill method and device, a storage medium, and an electronic device. Background Art
[0002] Traditional chaos engineering typically uses static fault injection strategies, which are unable to dynamically adjust the injection strategy based on the real-time system status. After fault injection, the system's response and recovery often require manual observation and analysis, without real-time feedback and experiment adjustments, resulting in low chaos experiment accuracy.
[0003] That is, there is a technical problem in the related technology that the accuracy of chaos engineering is low due to the use of static fault injection strategies.
[0004] Therefore, the problem of low accuracy of the chaos engineering fault drill results caused by performing chaos engineering fault drills through static fault injection strategies in related technologies has not yet been effectively solved. Summary of the Invention
[0005] The present application provides a chaos engineering fault drill method and device, a storage medium, and an electronic device to at least solve the problem of low accuracy of the chaos engineering fault drill results caused by performing chaos engineering fault drills through static fault injection strategies in related technologies.
[0006] The present application provides a chaos engineering fault drill method, comprising: an injection step: injecting an nth fault factor into a system to be tested, and obtaining an nth fault response result of the system to be tested under the action of the nth fault factor, so as to start a chaos engineering fault drill, where n is 1, 2, 3, etc. in sequence; a determination step: determining the current operating state of the system to be tested according to a target dimension and the nth fault response result, wherein the target dimension includes at least one of the following: a performance indicator dimension and a user experience dimension; an adjustment step: when it is determined that the current operating state is normal, adjusting the nth fault factor to an n+1th fault factor based on a fault adjustment rule and the nth fault response result; and cyclically executing the injection step, the determination step, and the adjustment step until an end time is reached, thereby terminating the chaos engineering fault drill for the system to be tested.
[0007] The present application also provides a chaos engineering fault drill device, including: an injection module, used to execute the injection step: inject the nth fault factor into the system to be tested, and obtain the nth fault response result of the system to be tested under the action of the nth fault factor, so as to start the chaos engineering fault drill, where n is 1, 2, 3... in sequence; a determination module, used to execute the determination step: determine the current operating status of the system to be tested according to the target dimension and the nth fault response result, wherein the target dimension includes at least one of the following: performance indicator dimension, user experience dimension; an adjustment module, used to execute the adjustment step: when it is determined that the current operating status is normal, adjust the nth fault factor to the n+1th fault factor based on the fault adjustment rule and the nth fault response result; a loop module, used to cyclically execute the injection step, the determination step and the adjustment step until the end time is reached, and end the chaos engineering fault drill for the system to be tested.
[0008] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned chaos engineering fault drill methods when executing the computer program.
[0009] The present application also provides a computer-readable storage medium, which stores a computer program, wherein when the computer program is executed by a processor, the steps of any of the above-mentioned chaos engineering fault drill methods are implemented.
[0010] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned chaos engineering fault drill methods when executed by a processor.
[0011] Through this application, since the nth fault factor is injected into the system under test and the nth fault response result of the system under test under the influence of the nth fault factor is obtained, the current operating state of the system under test is determined by the target dimension and the nth fault response result. When it is determined that the current operating state is normal, the nth fault factor is adjusted based on the fault adjustment rule and the nth fault response result, rather than adopting a static fault injection strategy. That is, the fault factor of this application is not fixed, but can be adjusted dynamically. Through the above technical solution, the problem of low accuracy of the chaotic engineering fault drill results caused by the static fault injection strategy in the related art can be solved, thereby improving the accuracy of chaos engineering. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0013] Figure 1 This is a hardware structure block diagram of a computer terminal for a chaos engineering fault drill method according to an embodiment of the present application;
[0014] Figure 2 This is a flowchart of a chaos engineering fault drill method according to an embodiment of the present application;
[0015] Figure 3 This is a flow chart of a method for performing data stream processing on a PCIE expansion card in the related art;
[0016] Figure 4 This is a flowchart of a method for real-time monitoring and dynamic adjustment of chaos engineering fault injection based on OpenTSDB according to an optional embodiment of the present application;
[0017] Figure 5 This is a structural block diagram of a chaos engineering fault drill device according to an embodiment of the present application. DETAILED DESCRIPTION
[0018] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0019] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0020] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0021] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the chaos engineering fault drill method depends, the specific application environment architecture or specific hardware architecture is described here.
[0022] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a computer terminal as an example, Figure 1 This is a hardware structure diagram of a computer terminal for a chaos engineering fault drill method according to an embodiment of the present application. Figure 1 As shown, the computer terminal may include one or more ( Figure 1 Only one is shown) a processor 102 (the processor 102 may include but is not limited to a microprocessor MCU or a programmable logic device FPGA and other processing devices) and a memory 104 for storing data. The computer terminal may also include a transmission device 106 and an input / output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above-mentioned computer terminal. For example, the computer terminal may also include Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0023] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the chaos engineering fault drill method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implementing the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0024] The transmission device 106 is used to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by a computer terminal's communications provider. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0025] The embodiments of this application provide a chaos engineering fault drill method. The following is an explanation of the technical terms involved in the embodiments of this application:
[0026] 1) Chaos Engineering is a cloud-native chaos engineering platform for multiple clusters, multiple environments, and multiple languages. It supports a wide range of experimental scenarios, including those involving the Central Processing Unit (CPU), memory, disk, and network. Chaos Engineering can simulate failures that may be encountered in actual production environments, and these failures are recoverable and will not affect the normal operation of the production environment.
[0027] 2) Open Time Series Database (OpenTSDB) is a distributed, scalable time series database that is often integrated into big data clusters. It is mainly used to store and manage time series data, providing efficient data access and statistical analysis capabilities.
[0028] 3) Prometheus (an open source monitoring system and time series database) is an open source monitoring software developed by Google. It can collect data from remote machines through the Hypertext Transfer Protocol (HTTP) and store it in a local time series database.
[0029] 4) Grafana (a monitoring visualization platform) is an open source monitoring instrumentation system that supports visualization using OpenTSDB as a data source.
[0030] Figure 2 This is a flowchart of a chaos engineering fault drill method according to an embodiment of the present application, which can be applied to Figure 1 In a computer terminal, such as Figure 2 As shown, the process includes the following steps:
[0031] Step S202, injection step: injecting the nth fault factor into the system under test and obtaining the nth fault response result of the system under test under the action of the nth fault factor to start the chaos engineering fault drill, where n is 1, 2, 3, etc.;
[0032] Step S204, determining step: determining the current operating state of the system under test according to the target dimension and the nth fault response result, wherein the target dimension includes at least one of the following: performance indicator dimension and user experience dimension;
[0033] Step S206, an adjustment step: when it is determined that the current operating state is a normal state, adjusting the nth fault factor to the n+1th fault factor based on a fault adjustment rule and the nth fault response result;
[0034] Step S208, looping through the injection step, the determination step, and the adjustment step until the end time is reached, thereby terminating the chaos engineering fault drill for the system under test.
[0035] Through the chaos engineering fault drill method of the present application, since the nth fault factor is injected into the system to be tested and the nth fault response result of the system to be tested under the action of the nth fault factor is obtained, the current operating state of the system to be tested is determined by the target dimension and the nth fault response result. When it is determined that the current operating state is normal, the nth fault factor is adjusted based on the fault adjustment rule and the nth fault response result, rather than adopting a static fault injection strategy. That is, the fault factor of the present application is not fixed, but can be dynamically adjusted. Through the above technical solution, the technical problem of low accuracy of chaos engineering caused by performing chaos engineering through static fault injection strategies in related technologies can be solved, thereby improving the accuracy of chaos engineering.
[0036] Optionally, in step S204, determining the current operating state of the system under test according to the dimensional indicators corresponding to the multiple target dimensions and the nth fault response result may include the following three situations:
[0037] (1) When the target dimension includes the performance indicator dimension, first performance indicator data corresponding to each performance indicator of the system under test during normal operation within a first time period is obtained; a baseline model corresponding to each performance indicator is established based on the first performance indicator data, and a baseline mean and standard deviation corresponding to the baseline model are calculated; second performance indicator data corresponding to each performance indicator of the system under test under the influence of the nth fault factor is obtained from the nth fault response result; a weight of each performance indicator in the system under test is determined based on a machine learning algorithm, and the product of the weight and the standard deviation is calculated; when it is determined that the second performance indicator data is greater than the sum of the baseline mean and the product, or the second performance indicator data is less than the difference between the baseline mean and the product, it is determined that the current operating state of the system under test is an abnormal state; when it is determined that the second performance indicator data is less than or equal to the sum of the baseline mean and the product, or the second performance indicator data is greater than or equal to the difference between the baseline mean and the product, it is determined that the current operating state of the system under test is a normal state.
[0038] It is understandable that the above performance indicators may include but are not limited to: latency (used to indicate request response time); throughput (used to indicate queries per second (QPS) / transactions per second (TPS)); resource utilization (used to indicate the use of CPU, memory, and disk input / output (I / O)), etc.
[0039] The above-mentioned baseline model is a statistical model used to describe the normal operating state of the system. In the embodiment of the present application, the baseline model is used to determine the normal threshold range of each performance indicator corresponding to the system to be tested.
[0040] In addition to the performance indicator dimension and the user experience dimension, the above-mentioned target dimensions may also include, for example: resource utilization dimension (analyzing the utilization rate of resources such as CPU, memory, disk space and network bandwidth), business logic dimension (analyzing the business logic execution of the system under test under fault conditions) and other dimensions. The embodiment of this application only takes the target dimensions including the performance indicator dimension and the user experience dimension as an example. Other dimensions are within the scope of protection of this application and are not listed one by one in this application.
[0041] If the target dimension includes a performance indicator dimension, first obtain the first performance indicator data corresponding to each performance indicator of the system under test during normal operation within a first time period. This first time period can be the period from when the system under test was added to Chaos Engineering but before Chaos Engineering fault drills began, for example, 30 days. A Prometheus monitoring agent can be deployed within Chaos Engineering and integrated with OpenTSDB. The Prometheus monitoring agent can then collect the first performance indicator data and store it in OpenTSDB to ensure efficient storage and querying of the first performance indicator data.
[0042] Furthermore, statistical algorithms such as mean can be used to establish a baseline model corresponding to each performance indicator based on the first performance indicator data, and the weight x of each performance indicator in the system to be tested can be calculated using a machine learning algorithm. In addition, the baseline mean and standard deviation of the baseline model can be obtained using the Prometheus function:
[0043] baseline_mean = avg_over_time(latency[nd]);
[0044] Standard deviation = stddev_over_time(latency[nd]).
[0045] In addition, it is also necessary to obtain second performance indicator data corresponding to each performance indicator of the system under test under the action of the nth fault response factor from the nth fault response result.
[0046] Then, the current operating status of the system under test is determined by the following formula:
[0047]
[0048] That is, when it is determined that the second performance indicator data is greater than the sum of the baseline mean and the product, or the second performance indicator data is less than the difference between the baseline mean and the product, the current operating state of the system under test is determined to be an abnormal state.
[0049] In addition, the above baseline model is not static, that is, the baseline model can be appropriately adjusted. Specifically:
[0050] Obtain each fault response result of the system under test under the action of each fault factor in the second time period, and obtain third performance indicator data corresponding to each performance indicator of the system under test under the action of each fault factor from each fault response result, wherein the third performance indicator data is the performance indicator data corresponding to each performance indicator when the operating state of the system under test is normal in the second time period; determine the periodic change law of the performance indicator data of each performance indicator of the system under test based on the first performance indicator data and the third performance indicator data, wherein the periodic change law includes at least one of the following: diurnal law, seasonal law; adjust the baseline model according to the periodic change law.
[0051] That is, the embodiment of the present application can make appropriate adjustments to the baseline model to adapt to the periodic changes in performance indicator data. Specifically:
[0052] During the chaos engineering fault drill on the system under test, the third performance indicator data of the system under test in a fault-free, normal operating state within the second time period can be collected. In the chaos engineering fault drill experiment, the chaos engineering fault drill experiment is divided into multiple time windows, and the second time period is one of the specific time periods, which is used to observe the system response of the system under test after the nth fault factor is injected.
[0053] Then, the time series of the first and third performance indicator data (both of which are performance indicator data for the system under test when it is operating normally) are analyzed, including but not limited to analysis of diurnal and seasonal patterns. Diurnal patterns refer to the periodic characteristics of system performance as it changes with time of day. For example, the system load during the day may be significantly higher than at night. Seasonal patterns reflect the cyclical patterns of system performance as it changes with the seasons over a longer period of time. For example, the usage of certain applications may increase significantly during holidays or certain months.
[0054] By collecting and analyzing the cyclical variations described above, we can adjust the baseline model to more accurately reflect the actual operating state of the system under test, particularly the periodic fluctuations that occur over time. This adjusted baseline model can better adapt to the varying operating states of the system under test during diurnal and seasonal variations. This allows for more accurate identification and response to abnormal behavior during chaos engineering drills, avoiding both false positives and false negatives.
[0055] (2) When the target dimension includes the user experience dimension, determine the fault type of the nth fault factor, and determine the target user experience indicator of the fault type among the multiple user experience indicators corresponding to the user experience dimension, wherein fault factors of different fault types correspond to different user experience indicators; determine the standard indicator range corresponding to the target user experience indicator, and obtain the indicator data indicated by the target user experience indicator of the system under test under the influence of the nth fault factor in the nth fault response result; when the indicator data is within the standard indicator range, determine that the current operating state of the system under test is normal; when the indicator data is not within the standard indicator range, determine that the current operating state of the system under test is abnormal.
[0056] It is understandable that the above target dimension may also include: user experience dimension. In the case where the above target dimension includes the user experience dimension:
[0057] Based on the nature of the nth failure factor (e.g., network latency, server downtime, database connection failure, etc.), select one or more sets of user experience metrics associated with it. User experience metrics may vary depending on the failure type. For example, a network latency failure may focus on metrics such as page load time and request response time, while a server downtime failure may focus more on service availability and response success rate.
[0058] Furthermore, a standard range of indicators under normal operating conditions is defined for each target user experience indicator. This standard range can be derived from statistical analysis of historical data or customized by the user, taking into account business characteristics and user expectations to ensure that the performance of the system under test under normal conditions meets user needs.
[0059] During the experiment, the target user experience indicator data of the system under test under the influence of the nth fault factor is obtained in real time. Then, the following judgment is made: If the monitored user experience indicator data falls within the standard indicator range, it indicates that the system under test can still maintain a good user experience under the fault condition, and the current system operation status is judged to be normal. Conversely, if the user experience indicator data exceeds the standard indicator range, it indicates that the fault has negatively affected the user experience, and the current status of the system under test is marked as abnormal.
[0060] By monitoring the user experience dimension in real time and combining it with the preset standard indicator range, we can provide instant feedback on the actual operating status of the system, and then dynamically adjust the fault injection strategy and the technical solution for triggering the recovery mechanism. This not only improves the accuracy and efficiency of chaos engineering experiments, but also effectively guarantees the service experience of end users.
[0061] (3) When the target dimension includes the performance indicator dimension and the user experience dimension, the evaluation results of the performance indicator dimension and the user experience dimension are considered together. If any target dimension is determined to be in an abnormal state, the overall evaluation determines that the current operating state of the system under test is in an abnormal state; only when all dimensions are determined to be in a normal state can the current operating state of the system under test be finally determined to be in a normal state.
[0062] Furthermore, based on the multi-dimensional evaluation results, the faults of the system under test can be graded, and the faults that have the greatest impact on system performance and user experience, such as critical service faults or core business logic errors, can be prioritized. A machine learning algorithm can be used to assess the impact of the fault on the system under test, and then determine the priority of the fault. The priority of the fault can also be customized by the user. After determining the priority of the fault, the highest priority fault is prioritized. After processing the highest priority fault, the first operating state of the system under test is determined. If the first operating state is determined to be an abnormal state, the faults with the next highest priority are processed in order of priority, and the first operating state of the system under test is determined again until the operating state of the system under test is determined to be normal.
[0063] Optionally, after determining the current operating state of the system under test according to the target dimension and the nth fault response result in the above step S204, the method further includes: when it is determined that the current operating state is an abnormal state, performing a fault recovery step on the system under test, wherein the fault recovery step includes at least one of the following: performing fault recovery on the system under test through a three-level recovery mechanism, the three-level recovery including: first-level recovery, second-level recovery and third-level recovery, the first-level recovery is to reduce the fault intensity corresponding to the nth fault factor, the second-level recovery is to stop the fault injection of the nth fault factor into the system under test, And perform a rollback operation on the system under test. The third-level recovery is to determine the target component in the system under test that has an abnormal state, isolate the target component from other components in the system under test except the target component, so as to prohibit the interaction between the target component and the other components, and trigger an alarm mechanism to instruct the target object to perform fault recovery on the target component and perform fault recovery on the system under test through a restart operation; when it is determined that the system under test in the abnormal state has been restored to a normal state, adjust the nth fault factor to the n+1th fault factor based on the fault adjustment rule and the nth fault response result.
[0064] It is understandable that, when the current operating state is determined to be an abnormal state, it is necessary to perform abnormal recovery on the system under test, and then perform the above adjustment steps on the system under test after abnormal recovery. The abnormal recovery method includes but is not limited to:
[0065] 1) Three-level recovery mechanism:
[0066] Level 1 recovery: attempts to reduce the intensity of the nth failure factor. Level 1 recovery is a relatively mild measure designed to minimize the impact on the system under test while observing whether the system under test can recover to normal state on its own.
[0067] Level 2 recovery: If level 1 recovery fails to restore the system under test to a normal state, the system will stop injecting the nth fault factor and perform a rollback, restoring the system to its most recent healthy state. Level 2 recovery is more aggressive than level 1 recovery and aims to quickly eliminate the impact of the failure.
[0068] Level 3 recovery: If level 2 recovery fails to resolve the system failure, it automatically identifies the target component in the system under test that is in an abnormal state, isolates it from other components in the system under test, prevents it from interacting with normal components, and triggers an alarm mechanism to notify the relevant operations and maintenance team (i.e., the target entity) to perform manual intervention and recovery. Level 3 recovery is the most stringent of the three levels of recovery, designed to prevent the spread of the anomaly and protect the system under test from further damage.
[0069] In the three-level recovery mechanism, after each step of the first level is completed, the key performance indicators and user experience indicators of the system under test must be continuously monitored to confirm whether the system under test has returned to normal. If the system under test has recovered successfully, further recovery measures will be stopped; otherwise, the recovery strategy will be continuously upgraded until the system under test returns to normal.
[0070] 2) Restarting the system to be tested to recover from the failure.
[0071] The three-level recovery and restart mechanisms in the above technical solution not only enable rapid response and resolution of abnormal situations encountered during chaos engineering fault drills, but also minimize the impact on the chaos engineering fault drill process and the system under test, ensuring the stable operation of the system under test.
[0072] Optionally, before injecting the nth fault factor into the system under test in the above step S202 and obtaining the nth fault response result of the system under test under the action of the nth fault factor, the method further includes: determining the business environment for performing the chaos engineering fault drill on the system under test and the business type corresponding to the chaos engineering fault drill, wherein the business environment includes at least one of the following: a production environment and a test environment, and the business type includes at least one of the following: a core business for indicating that a chaos engineering fault drill is performed on a core component in the system under test, and a non-core business for indicating that a chaos engineering fault drill is performed on a non-core component in the system under test; When it is determined that the business environment is the production environment, and / or the business type is the core business, the fault intensity of the nth fault factor is determined to be a fault intensity within a first fault intensity range, and the difference between the fault intensity of the nth fault factor and the fault intensity of the n+1th fault factor is determined to be less than or equal to a first fault intensity threshold; when it is determined that the business environment is the test environment, and the business type is the non-core business, the fault intensity of the nth fault factor is determined to be a fault intensity within a second fault intensity range, wherein the fault intensity within the second fault intensity range is greater than the fault intensity within the first fault intensity range.
[0073] It is understandable that the embodiment of the present application may also limit the fault intensity of the nth fault factor, specifically:
[0074] Determine the business environment and business type for the system under test to conduct chaos engineering fault drills:
[0075] When conducting chaos engineering fault drills in production environments, given the minimal direct impact on actual users and services, the fault intensity is limited to a first fault intensity range, and the difference in intensity between adjacent fault factors must be less than or equal to the first fault intensity threshold. Even if a fault injection causes a minor anomaly, rapid recovery is achieved, avoiding significant business impact.
[0076] In contrast, in a test environment, because the system under test is isolated from the actual production environment and has limited business impact, a wider range of fault intensity is tolerated, with the second fault intensity range exceeding the first. In a test environment, more aggressive fault injection strategies can be employed to more comprehensively examine the system under test's response to extreme conditions.
[0077] For core components in the system under test, such as databases and network services, since core components are directly related to the stability and business continuity of the system under test, even in the test environment, the intensity of the failure factor will be strictly controlled within the first failure intensity range to ensure that even if a failure occurs, it can be recovered quickly and effectively to avoid unacceptable damage to the core business.
[0078] For non-core components in the system under test, such as logging services and cache services, since these components have relatively little impact on the overall business, a higher level of fault injection within the second fault intensity range is permitted. This allows for a more in-depth test of system stability and the resilience of non-core components without impacting core business operations.
[0079] By dynamically adjusting the intensity of fault factors based on the business environment and type, we can maximize the stability and recovery capabilities of the system under test in various scenarios while ensuring the safety of chaos engineering fault drills.
[0080] Optionally, the above-mentioned step S206 of adjusting the nth fault factor to the n+1th fault factor based on the fault adjustment rule and the nth fault response result includes: determining the fault type corresponding to the nth fault factor, and determining the target performance indicator corresponding to the fault type in the fault adjustment rule; obtaining fourth performance indicator data corresponding to the target performance indicator of the system under test under the influence of the nth fault factor from the nth fault response result, and determining multiple fault adjustment conditions corresponding to the target performance indicator in the fault adjustment rule; determining the target fault adjustment condition reached by the nth fault factor according to the fourth performance indicator data, and determining the fault adjustment action of the nth fault factor under the target fault adjustment condition in the fault adjustment rule, wherein the multiple fault adjustment conditions include: the target fault adjustment condition; adjusting the nth fault factor to the n+1th fault factor according to the fault adjustment action.
[0081] It is understandable that the specific steps of adjusting the nth fault factor based on the fault rule include:
[0082] Determine the failure type and target performance metrics: During a chaos engineering failure drill, the first step is to identify the failure type of the nth failure factor, such as network latency, server downtime, or resource exhaustion. Based on the failure type, a set of closely related performance metrics is predefined. These metrics reflect the specific impact of the failure on the performance of the system under test, such as response time, throughput, and error rate.
[0083] Obtaining fourth performance indicator data and fault adjustment conditions: Obtain fourth performance indicator data corresponding to the target performance indicator of the system under test under the influence of the nth fault factor from the nth fault response result. Based on the fault type and target performance indicator, define a series of fault adjustment conditions. These conditions indicate whether the fault factor needs to be adjusted based on changes in the performance indicator. Fault adjustment conditions may include, but are not limited to, indicator data exceeding or falling below a certain threshold, and abnormal changes in the indicator trend.
[0084] Determine and execute fault adjustment actions: By comparing the current fourth performance indicator data with the preset fault adjustment conditions, determine whether one or more target fault adjustment conditions have been met. If a match is found, adjustment of the nth fault factor is necessary. Based on the preset fault adjustment rules, the corresponding fault adjustment action is searched for based on the achieved target fault adjustment conditions. Fault adjustment actions may include reducing fault severity, changing fault type, and increasing recovery resources. Based on these determinations and steps, the nth fault factor is adjusted to the n+1th fault factor.
[0085] Based on the above preset fault adjustment rules and conditions, fault management in chaos engineering fault drills is made more refined. Differentiated adjustment strategies can be adopted for different fault types and performance indicators, thereby improving the effectiveness and pertinence of chaos engineering fault drills.
[0086] In addition, since Grafana is integrated and OpenTSDB is configured during the chaos engineering fault drill, a dashboard can be created after the chaos engineering fault drill is completed, and the drill results corresponding to the chaos engineering fault drill can be displayed through the dashboard, wherein the drill results include at least one of the following: the relationship between the nth fault response result, the real-time change of the fault intensity of the nth fault factor and the nth fault response result; obtaining the historical drill results of the chaos engineering fault drill for the system to be tested, and determining the repair suggestions for the system to be tested based on the drill results and the historical drill results, wherein the repair suggestions include at least one of the following: short-term repair suggestions, long-term Repair suggestions; determining the degree of impact of each repair suggestion on the system under test, and calculating the repair cost corresponding to each repair suggestion; grading the repair suggestions based on the repair cost and degree of impact; generating a visualization report based on the drill results, wherein the visualization report includes at least one of the following: changes in performance indicator data corresponding to each performance indicator of the system under test under the influence of the nth and fault factors, abnormal behaviors of the system under test during the chaos engineering fault drill, abnormal recovery status of the system under test, the repair suggestions and the recommendation levels corresponding to the repair suggestions, and the abnormal behavior is used to indicate the behavior of the system under test when the current operating state is in an abnormal state.
[0087] Among them, the short-term repair suggestion can be, for example, expanding the machine capacity; the long-term repair suggestion can be, for example, optimizing the database index, etc.
[0088] Recommendation levels can include (A0-A2), and the repair recommendations for each level are as follows:
[0089] [A0] The threshold for the storage cluster network latency to enter maintenance mode has been adjusted from 75% to 80%.
[0090] Expected results: error rate decreased by 40%;
[0091] Recommendation basis: Historical similar case A;
[0092] [A1] Add one SSD to the storage cluster;
[0093] Expected effect: Storage performance improvement of 5%;
[0094] Basis for recommendation: Historical similar case B.
[0095] In order to better understand the process of the above-mentioned chaos engineering fault drill method, the implementation process of the above-mentioned chaos engineering fault drill method is described below in combination with an optional embodiment, but it is not used to limit the technical solution of the embodiment of this application.
[0096] With the advancement and development of technology, large-scale systems are becoming more and more common. Therefore, system stability has become particularly important, which has led to the emergence of chaos engineering. Chaos engineering simulates scenarios that may be encountered in production by injecting failures into the system, thereby identifying system weaknesses and making improvements.
[0097] Traditional chaos engineering typically uses static fault injection strategies, which are unable to dynamically adjust the injection strategy based on the real-time system status. After fault injection, the system's response and recovery often require manual observation and analysis, without real-time feedback and experiment adjustments. Furthermore, most tools lack the ability to model a baseline of normal system behavior, making it difficult to accurately identify abnormal behavior. After chaos experiments, the reported data is insufficiently detailed, making it difficult to intuitively understand the performance of key system performance parameters and the degree of system anomaly after the injection of a specific fault.
[0098] In order to improve the efficiency of chaos experiments and reduce the uncertainty caused by human intervention, the optional embodiment of this application proposes a method and device based on OpenTSDB real-time monitoring and dynamic adjustment of chaos engineering fault injection. Prometheus is used to collect data, and key system performance indicators are written to OpenTSDB in real time. Through statistical algorithm analysis, the fault injection strategy is dynamically adjusted, and the system is restored in time according to system performance. Real-time monitoring is combined with dynamic fault adjustment, timely fault detection and recovery to build a more intelligent chaos engineering, which comprehensively and accurately tests the stability and reliability of the system. Specifically:
[0099] 1. Device for real-time monitoring and dynamic adjustment of chaos engineering fault injection based on OpenTSDB:
[0100] Figure 3 This is a component architecture diagram of a device for real-time monitoring and dynamic adjustment of chaos engineering fault injection based on OpenTSDB according to an optional embodiment of the present application, such as Figure 3 As shown:
[0101] The device includes: a system to be tested, a monitoring agent module for data collection, OpenTSDB for data storage and query, a fault injection dynamic adjustment and recovery module for querying system real-time data and making dynamic adjustments, chaos engineering for fault drills (fault injection parameters need to be set within chaos engineering), and an indicator display module for displaying real-time data and reports.
[0102] The implementation of a device for real-time monitoring and dynamic adjustment of chaos engineering fault injection based on OpenTSDB mainly includes the following:
[0103] 1) Integrate OpenTSDB into the chaos engineering platform, leveraging its high-performance time series data storage capabilities to collect and store real-time metrics of the system under test (such as latency, throughput, and error rate), supporting multi-dimensional data analysis.
[0104] 2) Call the Prometheus monitoring agent to collect key system performance indicators in real time;
[0105] 3) Build a baseline model of normal system behavior based on key metrics stored in OpenTSDB;
[0106] 4) During the chaos experiment, the system performance data is monitored in real time and the fault injection parameters are adjusted dynamically.
[0107] 5) If abnormal system behavior is detected, the recovery mechanism is triggered;
[0108] 6) During the use of chaos, it supports visual viewing of changes in various indicators of the system under test;
[0109] 7) After the chaos experiment is completed, a detailed experimental report is output, including key indicators, fault injection effects, anomaly detection results and recovery status.
[0110] Specifically, a device that uses OpenTSDB for real-time monitoring and dynamic adjustment of chaos engineering fault injection can perform the following steps:
[0111] 1. Establish a baseline model.
[0112] (1) Deploy a monitoring agent and integrate OpenTSDB:
[0113] Deploy the Prometheus monitoring agent to collect key performance indicators of the system under test during normal operation and chaos experiments (i.e., chaos engineering fault drills). Store the key performance indicators of the system under test in OpenTSDB to ensure efficient storage and query of data.
[0114] (2) Configure a chaos experiment (taking the example of injecting network delay failures into a distributed storage system to verify system stability):
[0115] Create a chaos engineering drill, define fault injection (i.e., fault factor) as injecting network delay into the specified distributed storage cluster, and configure the initial intensity of the network delay to 50ms.
[0116] (3) Establish a baseline model of normal system behavior:
[0117] By default, the system's performance indicators (including latency, throughput, error rate, CPU, memory, and disk usage) for the past 30 days are extracted from OpenTSDB. If the system has been added to Chaos Engineering for less than 30 days, the maximum number of days is used for extraction. Custom settings are also supported, with the number of days being recorded as n (i.e., the first time period).
[0118] Use statistical algorithms such as mean to establish a baseline model, calculate x (i.e., weight) through machine learning algorithms, and determine the normal threshold range of the system (i.e., the method of judging the current operating status of the system under test by the performance indicator dimension):
[0119]
[0120] In addition, the current operating status of the system under test can also be judged from the user experience dimension:
[0121] In the user experience dimension, when the system (i.e., the system under test) is performing fault adjustment, while comparing the baseline model indicators, it also considers whether some indicators that affect the user experience exceed a certain value.
[0122] User experience metrics for the user experience dimension include interactivity (page load time), functionality (function error frequency), and continuity (task completion rate). Specific metric thresholds can be customized. Different metrics are considered depending on the injected fault.
[0123] When performing dynamic fault adjustments, the system performance and user experience dimensions work together. For example, when network fault latency is injected into the system, during the drill, system performance metrics are acquired in real time. Using the aforementioned criteria, the system determines whether an anomaly exists. Simultaneously, user experience metrics are automatically selected based on the fault type for further evaluation. Network fault latency corresponds to interactivity metrics, and whether these metrics are met is determined. If not, adjustments or recovery are triggered.
[0124] In an optional embodiment of this application, a baseline corresponding to a real-time baseline model can be calculated based on a rolling time window (e.g., the past 5 minutes) to quickly respond to sudden anomalies. Alternatively, OpenTSDB can be used to store historical data and analyze diurnal / seasonal patterns (i.e., periodic patterns) to optimize the baseline model.
[0125] At the same time, the fault type of the injected fault can also be determined. Different fault types can correspond to different adjustment strategies. For example, Table 1 shows the adjustment strategies corresponding to the fault types according to an optional embodiment of the present application, as shown in Table 1:
[0126] Table 1
[0127] Fault type Adjustment strategy Network latency Dynamically adjust intensity based on latency deviation from baseline Pod Kill Adjust trigger frequency based on service recovery time CPU Load Combine CPU utilization with business traffic to make decisions
[0128] And it can combine different fault types to implement a brake priority mechanism: for example, critical service faults (such as database downtime) prioritize downgrading the injection intensity.
[0129] In addition, optional embodiments of this application can also determine fault injection strategies based on business scenarios. For example, when testing in a production environment, a conservative measurement is adopted by default, covering 99% of policy data to control risks; when verifying in a test environment, an aggressive strategy is adopted to quickly expose potential risks. When testing in a business type, for example, the frequency or intensity of fault injection is automatically reduced for core businesses such as databases, prioritizing stability; for non-core modules such as logging, high-intensity fault injection is enabled.
[0130] Optional embodiments of this application also provide some risk control measures: setting safety thresholds: setting a fuse mechanism (such as stopping the experiment if the error rate is >10% for three consecutive times). Gradual adjustment: increasing / decreasing the fault intensity in small steps (such as ±10%) to avoid system crashes.
[0131] 2. Perform fault injection and fault adjustment according to the fault adjustment rules:
[0132] 1) Before the experiment begins, the device supports dragging and dropping indicators, conditions, and actions to define dynamic adjustment rules (i.e., fault adjustment rules). For example, Table 2 shows a schematic diagram of fault adjustment rules when the fault injection is determined to be delay, as shown in Table 2:
[0133] Table 2
[0134] index condition action Fault type Rule 1 System latency >120ms reduce Network latency Rule 2 System latency <80ms Increase Network latency
[0135] 2) During the chaos drill, the device will query real-time system performance data from OpenTSDB every 10 seconds, and based on pre-defined dynamic adjustment rules, it will determine the fault parameters that need to be dynamically adjusted. It will then call the Chaos Engineering API to modify the fault injection parameters. For example, if the current fault injection time is 50ms and the system delay is detected to be <100ms, the network delay fault injection will be increased to 70ms; if the system delay is detected to be >120ms, the fault injection will be reduced to 60ms.
[0136] 3) After dynamically adjusting the fault injection, the adjusted fault parameters and system performance indicators will be fed back to OpenTSDB in real time, forming a closed-loop control.
[0137] 4) If a system anomaly is detected, the recovery mechanism is triggered:
[0138] During a chaos experiment, current metrics are compared with the baseline model. Based on the upper and lower thresholds set by the baseline model, any deviations from these thresholds are considered abnormal. For example, if the baseline latency is calculated to have a mean of 100ms and a standard deviation of 20ms, and x = 2, then a system abnormality is considered if the system network latency is >140ms or <60ms. This technical solution allows for dynamic baseline recovery, eliminating the need for manual threshold setting. This approach is more responsive to actual conditions and can adapt to dynamic system fluctuations.
[0139] If a system anomaly indicator is detected, fault injection is stopped and recovery is performed according to pre-set recovery methods, such as restarting the storage cluster. Abnormal recovery mechanisms can include, for example: built-in recovery strategies: 1) a three-level recovery mechanism with progressive recovery capabilities, including degradation (reducing the severity of the fault), rollback (stopping the current fault injection and restoring the system under test to its most recent healthy state), and isolation (isolating the faulty module and triggering a critical alarm); 2) custom recovery methods (supporting pre-set recovery methods, such as restarting the cluster).
[0140] After the recovery strategy is triggered, the system continuously monitors the indicators until they return to the baseline range (relying on the millisecond-level response of OpenTSDB).
[0141] 3. Collection and display of experimental results:
[0142] Supports viewing of real-time system performance indicators and generates visual test reports after the drill: Integrate Grafana and configure the data source to be OpenTSDB; create a dashboard to display the system's key performance indicators and fault injection effects during the drill, and show the real-time relationship between latency data and fault injection intensity.
[0143] After the drill, a visual report is generated. The report content includes: experiment overview (fault injection type, test duration and other basic information), changes in system performance indicators, abnormal behavior, system recovery status, conclusions and suggestions (based on the current experimental data, conclusions are given, and combined with historical experimental records, intelligent repair suggestions are given, including: short-term suggestions: such as expanding the machine capacity; long-term suggestions: such as optimizing database indexes), and the given suggestions are graded (A0-A5) according to the degree of impact and repair cost.
[0144] Repair suggestions such as:
[0145] [A0] The threshold for the storage cluster network latency to enter maintenance mode has been adjusted from 75% to 80%.
[0146] Expected results: error rate decreased by 40%;
[0147] Recommendation basis: Historical similar case A;
[0148] [A1] Add one SSD to the storage cluster;
[0149] Expected effect: Storage performance improvement of 5%;
[0150] Basis for recommendation: Historical similar case B.
[0151] 2. Real-time monitoring and dynamic adjustment of chaos engineering fault injection based on OpenTSDB:
[0152] Figure 4 This is a flowchart of a method for real-time monitoring and dynamic adjustment of chaos engineering fault injection based on OpenTSDB according to an optional embodiment of the present application. Figure 4 As shown:
[0153] Step S401, creating a chaos drill;
[0154] You need to configure the system information to be tested, set the fault injection type and initial fault injection parameters, define the fault adjustment strategy, and set the normal behavior baseline value number n, where n can be 30 days.
[0155] Step S402: Obtain the time D when the system under test is added to the chaos project;
[0156] Step S403, determine whether D ≥ n;
[0157] If it is determined that D≥n, step S404 is executed; if it is determined that D<n, step S405 is executed;
[0158] Step S404, when it is determined that D≥n, the number of days n of the normal behavior baseline value is taken as n;
[0159] Step S405: If it is determined that D is less than n, the number of days n of the normal behavior baseline value is taken as D;
[0160] Step S406: establishing a normal behavior baseline (i.e., a baseline model) based on the performance data of the previous n days;
[0161] Step S407, starting the chaos experiment;
[0162] You need to view system performance indicators in real time and store them in OpenTSDB to view changes in system performance indicators in real time.
[0163] Step S408: Obtain performance data stored in OpenTSDB every 10 seconds.
[0164] Step S409, determining whether the drill end time has been reached;
[0165] If it is determined that the drill end time has arrived, step S413 is executed;
[0166] If it is determined that the drill end time has not been reached, step S410 is executed;
[0167] Step S410, determining whether the system has abnormal behavior;
[0168] If it is determined that the system has abnormal behavior, step S411 is executed; if it is determined that the system has no abnormal behavior, step S412 is executed;
[0169] Step S411, triggering an automatic recovery mechanism when it is determined that the system has abnormal behavior;
[0170] Step S412: When it is determined that there is no abnormal behavior in the system, the fault injection strategy is dynamically adjusted.
[0171] Step S413, when it is determined that the drill end time has been reached, end the drill;
[0172] The above steps S408 to S412 are executed in a loop until the drill end time is reached, and the drill ends.
[0173] Step S414: Generate a visualization report.
[0174] In summary, the optional embodiment of the present application is based on OpenTSDB, which implements real-time monitoring and dynamic adjustment of fault injection during chaos engineering fault drills, thereby building a more efficient chaos engineering device and method. Through the above-mentioned device and method, the intensity, type and scope of fault injection are dynamically adjusted according to the real-time monitoring data in OpenTSDB, making the experiment closer to the actual production environment. Through automated fault detection and recovery, manual intervention is reduced and efficiency is improved. A baseline model is established to automatically detect abnormal behavior, avoid system crashes due to excessive fault injection, or fail to discover potential problems due to insufficient injection. After an anomaly is detected, a recovery mechanism is triggered to shorten system recovery time and reduce the impact on the business. In addition, through OpenTSDB and Grafana, experimental data is displayed in real time to provide intuitive experimental results, and after the experiment, a detailed report containing key indicators, anomaly detection results and recovery status is generated.
[0175] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.
[0176] This embodiment also provides a chaos engineering fault drill device for implementing the aforementioned embodiments and preferred implementations. Details already described are omitted. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware, is also possible and contemplated.
[0177] Figure 5 This is a structural block diagram of a chaos engineering fault drill device according to an embodiment of the present application. Figure 5 As shown, the device includes:
[0178] An injection module 52 is configured to perform an injection step: inject an nth fault factor into the system under test, and obtain an nth fault response result of the system under test under the action of the nth fault factor, so as to start a chaos engineering fault drill, where n is 1, 2, 3, etc.;
[0179] The determining module 54 is configured to perform a determining step: determining a current operating state of the system under test according to a target dimension and the nth fault response result, wherein the target dimension includes at least one of the following: a performance indicator dimension and a user experience dimension;
[0180] an adjustment module 56 configured to perform an adjustment step: when determining that the current operating state is a normal state, adjusting the nth fault factor to an n+1th fault factor based on a fault adjustment rule and the nth fault response result;
[0181] The loop module 58 is used to cyclically execute the injection step, the determination step and the adjustment step until the end time is reached, thereby ending the chaos engineering fault drill on the system under test.
[0182] Through the chaos engineering fault drill device of the present application, since the nth fault factor is injected into the system to be tested and the nth fault response result of the system to be tested under the action of the nth fault factor is obtained, the current operating state of the system to be tested is determined by the target dimension and the nth fault response result. When it is determined that the current operating state is normal, the nth fault factor is adjusted based on the fault adjustment rule and the nth fault response result, instead of adopting a static fault injection strategy. That is, the fault factor of the present application is not fixed, but can be adjusted dynamically. Through the above technical solution, the technical problem of low accuracy of chaos engineering caused by performing chaos engineering through static fault injection strategy in related technologies can be solved, thereby improving the accuracy of chaos engineering.
[0183] In an exemplary embodiment, the determination module 54 is further configured to, when the target dimension includes the performance indicator dimension, obtain first performance indicator data corresponding to each performance indicator of the system under test during normal operation within a first time period; establish a baseline model corresponding to each performance indicator based on the first performance indicator data, and calculate a baseline mean and standard deviation corresponding to the baseline model; obtain second performance indicator data corresponding to each performance indicator of the system under test under the influence of the nth fault factor from the nth fault response result; determine a weight of each performance indicator in the system under test based on a machine learning algorithm, and calculate the product of the weight and the standard deviation; determine that the current operating state of the system under test is an abnormal state if it is determined that the second performance indicator data is greater than the sum of the baseline mean and the product, or that the second performance indicator data is less than the difference between the baseline mean and the product; and determine that the current operating state of the system under test is a normal state if it is determined that the second performance indicator data is less than or equal to the sum of the baseline mean and the product, or that the second performance indicator data is greater than or equal to the difference between the baseline mean and the product.
[0184] In an exemplary embodiment, the determination module 54 is further used to determine the fault type of the nth fault factor when the target dimension includes the user experience dimension, and determine the target user experience indicator of the fault type among the multiple user experience indicators corresponding to the user experience dimension, wherein fault factors of different fault types correspond to different user experience indicators; determine the standard indicator range corresponding to the target user experience indicator, and obtain the indicator data indicated by the target user experience indicator of the system under test under the influence of the nth fault factor in the nth fault response result; when the indicator data is within the standard indicator range, determine that the current operating state of the system under test is normal; when the indicator data is not within the standard indicator range, determine that the current operating state of the system under test is abnormal.
[0185] In an exemplary embodiment, the determination module 54 is also used to obtain each fault response result of the system under test under the influence of each fault factor in the second time period, and obtain third performance indicator data corresponding to each performance indicator of the system under test under the influence of each fault factor from each fault response result, wherein the third performance indicator data is the performance indicator data corresponding to each performance indicator when the operating state of the system under test is normal in the second time period; determine the periodic change law of the performance indicator data of each performance indicator of the system under test based on the first performance indicator data and the third performance indicator data, wherein the periodic change law includes at least one of the following: diurnal law, seasonal law; adjust the baseline model according to the periodic change law.
[0186] In an exemplary embodiment, the adjustment module 56 is further configured to, upon determining that the current operating state is an abnormal state, perform a fault recovery step on the system under test, wherein the fault recovery step includes at least one of the following: performing fault recovery on the system under test through a three-level recovery mechanism, wherein the three-level recovery includes: a first-level recovery, a second-level recovery, and a third-level recovery, wherein the first-level recovery is to reduce the fault intensity corresponding to the n-th fault factor, the second-level recovery is to stop fault injection of the n-th fault factor into the system under test and perform a rollback operation on the system under test, and the third-level recovery is to determine a target component in the system under test in which an abnormal state occurs, isolate the target component from other components in the system under test except the target component to prohibit interaction between the target component and the other components, and trigger an alarm mechanism to instruct the target object to perform fault recovery on the target component and perform fault recovery on the system under test through a restart operation; and upon determining that the system under test in the abnormal state has recovered to a normal state, adjusting the n-th fault factor to the n+1-th fault factor based on the fault adjustment rule and the n-th fault response result.
[0187] In an exemplary embodiment, the adjustment module 56 is further used to determine the business environment for performing the chaos engineering fault drill on the system under test and the business type corresponding to the chaos engineering fault drill, wherein the business environment includes at least one of the following: a production environment and a test environment, and the business type includes at least one of the following: a core business for indicating that a chaos engineering fault drill is performed on the core components of the system under test, and a non-core business for indicating that a chaos engineering fault drill is performed on the non-core components of the system under test; when it is determined that the business environment is the production environment and / or the business type is the core business, it is determined that the fault intensity of the nth fault factor is a fault intensity within a first fault intensity range, and it is determined that the difference between the fault intensity of the nth fault factor and the fault intensity of the (n+1)th fault factor is less than or equal to a first fault intensity threshold; when it is determined that the business environment is the test environment and the business type is the non-core business, it is determined that the fault intensity of the nth fault factor is a fault intensity within a second fault intensity range, wherein the fault intensity within the second fault intensity range is greater than the fault intensity within the first fault intensity range.
[0188] In an exemplary embodiment, the adjustment module 56 is further configured to determine the fault type corresponding to the nth fault factor, and determine the target performance indicator corresponding to the fault type in the fault adjustment rule; obtain fourth performance indicator data corresponding to the target performance indicator of the system under test under the influence of the nth fault factor from the nth fault response result, and determine multiple fault adjustment conditions corresponding to the target performance indicator in the fault adjustment rule; determine the target fault adjustment condition reached by the nth fault factor based on the fourth performance indicator data, and determine the fault adjustment action of the nth fault factor under the target fault adjustment condition in the fault adjustment rule, wherein the multiple fault adjustment conditions include: the target fault adjustment condition; and adjusting the nth fault factor to the n+1th fault factor according to the fault adjustment action.
[0189] For descriptions of the features in the embodiments corresponding to the chaos engineering fault drill apparatus, please refer to the relevant descriptions of the embodiments corresponding to the chaos engineering fault drill method, and will not be repeated here.
[0190] An embodiment of the present application also provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned chaos engineering fault drill method embodiments.
[0191] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned chaos engineering fault drill method embodiments when running.
[0192] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0193] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned chaos engineering fault drill method embodiments are implemented.
[0194] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps in any of the above-mentioned chaos engineering fault drill method embodiments.
[0195] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0196] The above is a detailed introduction to a chaos engineering fault drill provided by this application. This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the scope of protection of the claims of this application.
Claims
1. A chaos engineering fault drill method, characterized in that: include: Injection step: Inject the nth fault factor into the system under test and obtain the nth fault response result of the system under test under the influence of the nth fault factor to start the chaos engineering fault drill. n is 1, 2, 3, etc. Determining step: determining the current operating state of the system under test according to a target dimension and the nth fault response result, wherein the target dimension includes at least one of the following: a performance indicator dimension and a user experience dimension; Adjustment step: when it is determined that the current operating state is a normal state, adjusting the nth fault factor to the n+1th fault factor based on a fault adjustment rule and the nth fault response result; The injection step, the determination step, and the adjustment step are executed in a loop until the end time is reached, thereby completing the chaos engineering fault drill on the system under test.
2. The chaos engineering fault drill method according to claim 1, characterized in that: Determining a current operating state of the system under test according to the target dimension and the nth fault response result includes: In a case where the target dimension includes the performance indicator dimension, obtaining first performance indicator data corresponding to each performance indicator of the system under test during normal operation within a first time period; Establishing a baseline model corresponding to each performance indicator according to the first performance indicator data, and calculating a baseline mean and a standard deviation corresponding to the baseline model; Obtaining second performance indicator data corresponding to each performance indicator of the system under test under the influence of the nth fault factor from the nth fault response result; Determine the weight of each performance indicator in the system to be tested according to a machine learning algorithm, and calculate the product of the weight and the standard deviation; When it is determined that the second performance indicator data is greater than the sum of the baseline mean and the product, or the second performance indicator data is less than the difference between the baseline mean and the product, determining that the current operating state of the system under test is an abnormal state; When it is determined that the second performance indicator data is less than or equal to the sum of the baseline mean and the product, or the second performance indicator data is greater than or equal to the difference between the baseline mean and the product, it is determined that the current operating state of the system under test is normal.
3. The chaos engineering fault drill method according to claim 1, characterized in that: Determining a current operating state of the system under test according to the target dimension and the nth fault response result includes: In a case where the target dimension includes the user experience dimension, determining a fault type of the nth fault factor, and determining a target user experience index from a plurality of user experience indexes corresponding to the fault type, wherein fault factors of different fault types correspond to different user experience indexes; Determining a standard indicator range corresponding to the target user experience indicator, and obtaining indicator data indicated by the target user experience indicator of the system under test under the influence of the nth fault factor in the nth fault response result; When the indicator data is within the standard indicator range, determining that the current operating state of the system to be tested is a normal state; When the indicator data is not within the standard indicator range, it is determined that the current operating state of the system to be tested is an abnormal state.
4. The chaos engineering fault drill method according to claim 2, characterized in that: After establishing the baseline model corresponding to each performance indicator according to the first performance indicator data, the method further includes: Obtaining each fault response result of the system under test under the action of each fault factor in the second time period, and obtaining third performance indicator data corresponding to each performance indicator of the system under test under the action of each fault factor from each fault response result, wherein the third performance indicator data is performance indicator data corresponding to each performance indicator when the operating state of the system under test in the second time period is normal; Determining a periodic variation law of performance indicator data of each performance indicator of the system to be tested based on the first performance indicator data and the third performance indicator data, wherein the periodic variation law includes at least one of the following: a diurnal law and a seasonal law; The baseline model is adjusted according to the periodic variation rule.
5. The chaos engineering fault drill method according to claim 1, characterized in that: After determining the current operating state of the system under test according to the target dimension and the nth fault response result, the method further includes: In the case where it is determined that the current operating state is an abnormal state, a fault recovery step is performed on the system under test, wherein the fault recovery step includes at least one of the following: performing fault recovery on the system under test through a three-level recovery mechanism, the three-level recovery mechanism including: first-level recovery, second-level recovery and third-level recovery, the first-level recovery is to reduce the fault intensity corresponding to the n-th fault factor, the second-level recovery is to stop the fault injection of the n-th fault factor into the system under test and perform a rollback operation on the system under test, the third-level recovery is to determine a target component in the system under test that has an abnormal state, isolate the target component from other components in the system under test except the target component to prohibit interaction between the target component and the other components, and trigger an alarm mechanism to instruct the target object to perform fault recovery on the target component and perform fault recovery on the system under test through a restart operation; When it is determined that the system under test in the abnormal state has recovered to a normal state, the nth fault factor is adjusted to the (n+1)th fault factor based on the fault adjustment rule and the nth fault response result.
6. The chaos engineering fault drill method according to claim 1, characterized in that: Before injecting the nth fault factor into the system under test and obtaining the nth fault response result of the system under test under the action of the nth fault factor, the method further includes: Determine the business environment for performing the chaos engineering fault drill on the system under test and the business type corresponding to the chaos engineering fault drill, wherein the business environment includes at least one of the following: a production environment and a test environment, and the business type includes at least one of the following: a core business for indicating that a chaos engineering fault drill is performed on a core component in the system under test, and a non-core business for indicating that a chaos engineering fault drill is performed on a non-core component in the system under test; When it is determined that the business environment is the production environment and / or the business type is the core business, determining that the fault intensity of the nth fault factor is a fault intensity within a first fault intensity range, and determining that a difference between the fault intensity of the nth fault factor and the fault intensity of the (n+1)th fault factor is less than or equal to a first fault intensity threshold; When it is determined that the business environment is the test environment and the business type is the non-core business, the fault intensity of the nth fault factor is determined to be a fault intensity within a second fault intensity range, wherein the fault intensity within the second fault intensity range is greater than the fault intensity within the first fault intensity range.
7. The chaos engineering fault drill method according to claim 1, characterized in that: Adjusting the nth fault factor to an n+1th fault factor based on a fault adjustment rule and the nth fault response result includes: Determining a fault type corresponding to the nth fault factor, and determining a target performance indicator corresponding to the fault type in the fault adjustment rule; Obtaining fourth performance indicator data corresponding to the target performance indicator of the system under test under the influence of the nth fault factor from the nth fault response result, and determining a plurality of fault adjustment conditions corresponding to the target performance indicator in the fault adjustment rule; determining a target fault adjustment condition reached by the nth fault factor based on the fourth performance indicator data, and determining a fault adjustment action for the nth fault factor under the target fault adjustment condition in the fault adjustment rule, wherein the plurality of fault adjustment conditions include: the target fault adjustment condition; The nth fault factor is adjusted to the n+1th fault factor according to the fault adjustment action.
8. A chaos engineering fault drill device, characterized in that: include: An injection module is configured to perform an injection step: inject an nth fault factor into the system under test and obtain an nth fault response result of the system under test under the action of the nth fault factor to start a chaos engineering fault drill, where n is 1, 2, 3, etc.; a determination module, configured to execute a determination step: determining a current operating state of the system under test according to a target dimension and the nth fault response result, wherein the target dimension includes at least one of the following: a performance indicator dimension and a user experience dimension; an adjustment module, configured to perform an adjustment step: when determining that the current operating state is a normal state, adjusting the nth fault factor to an n+1th fault factor based on a fault adjustment rule and the nth fault response result; A loop module is used to cyclically execute the injection step, the determination step and the adjustment step until the end time is reached, thereby ending the chaos engineering fault drill on the system under test.
9. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the chaos engineering fault drill method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the chaos engineering fault drill method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Computing hardware system testing method based on chaos engineering
CN120929346A