A Resilience Measurement Method and Device for Distributed Systems Based on Chaos Engineering

Through the chaos engineering platform, the problem of low manual operation efficiency in the existing technology is solved, and the comprehensiveness and intuitiveness of distributed system toughness measurement is achieved, and the efficiency and accuracy of testing are improved.

CN117762669BActive Publication Date: 2025-07-18CHINA ELECTRONICS JINXIN DIGITAL TECH GRP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311772979.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-20
Publication Date
2025-07-18
Estimated Expiration
2043-12-20

AI Technical Summary

Technical Problem

In the prior art, distributed system toughness measurement methods that rely on manual operations are inefficient and have high probability of errors, and the measurement results are difficult to intuitively perceive, so the system toughness cannot be comprehensively evaluated.

Method used

Using a method based on chaos engineering, a user interaction interface is provided through the chaos engineering platform, fault scenarios, pressure scenarios and data items to be monitored are configured, fault drills are automatically performed, and test data is collected, and the toughness test results are determined based on the test results, which are displayed in the toughness metric table.

Benefits of technology

It realizes the completeness and efficiency improvement of distributed system toughness measurement, can intuitively present test results and progress, reduce manual operations, and improve the degree of automation and accuracy of tests.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117762669B_ABST
    Figure CN117762669B_ABST
Patent Text Reader

Abstract

The present application provides a method and device for measuring the resilience of a distributed system based on chaos engineering, which are applied to a chaos engineering platform. The chaos engineering platform provides a user interaction interface; the user interaction interface displays a resilience measurement table of the target distributed system; the method includes: for each test item in the resilience measurement table, in response to a configuration operation on the chaos engineering platform, configuring a failure scenario, a stress scenario, and a data item to be monitored corresponding to the test item in the resilience measurement of the target distributed system; performing a failure drill on the target distributed system according to the failure scenario and the stress scenario, and collecting test data corresponding to the data item to be monitored; determining a resilience test result under the test item according to the test data; and displaying the resilience test result under the test item at the corresponding display position in the resilience measurement table. In this way, the test results and the test progress can be visually presented, and the integrity and efficiency of measuring the resilience of the distributed system can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of distributed systems, and in particular, to a method and device for measuring the resilience of a distributed system based on chaos engineering. Background Art

[0002] To measure the resilience of a distributed system, the prior art generally simulates faults through manual command typing or program script execution for system testing. However, this manual operation-dependent method is inefficient, has a high error probability, provides an incomplete measurement of the resilience of the distributed system, and the measurement results are difficult to intuitively perceive. Summary of the Invention

[0003] In view of this, the purpose of this application is to provide a method and device for measuring the resilience of a distributed system based on chaos engineering to solve the problems in the prior art that the manual operation-dependent testing method is inefficient, has a high error probability, provides an incomplete measurement of the resilience of the distributed system, and the measurement results are difficult to intuitively perceive.

[0004] This application also provides a method for measuring the resilience of a distributed system based on chaos engineering, which is applied to a chaos engineering platform that provides a user interface; the user interface displays a resilience measurement table of the target distributed system; the method includes:

[0005] For each test item in the resilience measurement table, in response to a configuration operation on the chaos engineering platform, configure the fault scenario, stress scenario, and data items to be monitored corresponding to this test item in the resilience measurement of the target distributed system;

[0006] Execute a fault drill on the target distributed system according to the fault scenario and the stress scenario, and collect the test data corresponding to the data items to be monitored;

[0007] Determine the resilience test result of the target distributed system under this test item according to the test data;

[0008] Display the resilience test result under this test item at the display position corresponding to this test item in the resilience measurement table.

[0009] Further, for each test item in the resilience measurement table, in response to a configuration operation on the chaos engineering platform, configuring the fault scenario, stress scenario, and data items to be monitored corresponding to this test item in the resilience measurement of the target distributed system includes:

[0010] For each test item in the resilience measurement scale, in response to a first configuration operation on the chaos tool in the chaos engineering platform, select the types of fault events under this test item, and configure the corresponding fault drill objects and fault parameters in the target distributed system for each selected fault event, generating a fault scenario for each fault event;

[0011] In response to a second configuration operation on the test tool in the chaos engineering platform, configure the stress parameters for each fault event, and generate a stress scenario for each fault event according to the stress parameters and the fault drill objects;

[0012] In response to a third configuration operation on the system monitoring platform in the chaos engineering platform, determine the resource data and business metrics related to each fault event as the data items to be monitored; and in response to a fourth configuration operation on the service registration and discovery component in the chaos engineering platform, determine the service instance status parameters of the fault drill objects as the data items to be monitored.

[0013] Furthermore, perform a fault drill on the target distributed system according to the fault scenario and the stress scenario, and collect the test data corresponding to the data items to be monitored, including:

[0014] For each fault event under this test item, the test tool sends a data stream to the fault drill object according to the stress parameters corresponding to the stress scenario to apply stress;

[0015] During the stress application process, the chaos tool sends a fault injection instruction to the fault drill object according to the fault parameters corresponding to the fault scenario, so that the Agent probe pre-installed in the target distributed system executes fault event injection on the fault drill object according to the fault injection instruction;

[0016] The system monitoring platform and the service registration and discovery component collect the test data corresponding to the data items to be monitored of the fault drill object.

[0017] Furthermore, determine the resilience test result of the target distributed system under this test item according to the test data, including:

[0018] For each fault event under this test item, determine the sub-resilience test result of the target distributed system under this fault event according to the test data and the steady-state hypothesis configured for this fault event in the resilience measurement scale;

[0019] Integrate the sub-resilience test results under each fault event to determine the resilience test result of the target distributed system under this test item.

[0020] Further, for each failure event under the test item, according to the test data and the steady-state assumptions configured for this failure event in the resilience measurement table, determine the sub-resilience test result of the target distributed system under this failure event, including:

[0021] For each failure event under the test item, analyze the steady-state assumptions configured for this failure event, and parse out the target parameters constrained by the steady-state assumptions and the numerical judgment conditions of the target parameters;

[0022] Determine the target data items to be monitored corresponding to the target parameters and the calculation relationship between the target data items to be monitored;

[0023] Extract the target test data corresponding to the target data items to be monitored from the test data, and obtain the parameter values of the target parameters according to the calculation relationship;

[0024] According to the parameter values of the target parameters and the numerical judgment conditions of the target parameters, determine the sub-resilience test result of the target distributed system under this failure event.

[0025] Further, construct the resilience measurement table of the target distributed system in the following manner:

[0026] In response to the import operation for the chaos engineering platform, import the pre-constructed original resilience measurement table into the chaos engineering platform and display it on the user interface;

[0027] In response to the selection operation for the original resilience measurement table in the chaos engineering platform, select the test index dimensions, the index sub-dimensions under each test index dimension, and the test items under each index sub-dimension from the original resilience measurement table to construct the resilience measurement table of the target distributed system.

[0028] Further, the resilience test results include all passed, all not passed, and partially passed; the method further includes:

[0029] Mark this test item as a completed test item;

[0030] According to the proportion of the completed test items in the resilience measurement table, determine the resilience measurement progress of the target distributed system, and display the resilience measurement progress at the display position corresponding to the resilience measurement progress item in the resilience measurement table;

[0031] According to the test items in the resilience measurement table with the resilience measurement results of all not passed and partially passed, determine the weak test index dimensions of the target distributed system.

[0032] Further, after displaying the resilience test result under the test item at the corresponding display position of the resilience measurement table, the method further includes:

[0033] lifting the fault drill performed on the target distributed system;

[0034] determining the system recovery situation of the target distributed system according to the system data corresponding to the data item to be monitored;

[0035] After the system recovery situation indicates that the target distributed system has returned to normal, execute the resilience measurement method for the next test item.

[0036] An embodiment of the present application further provides a resilience measurement device for a distributed system based on chaos engineering, which is applied to a chaos engineering platform that provides a user interaction interface; the user interaction interface displays a resilience measurement table of a target distributed system; the device includes:

[0037] A configuration module, configured to, for each test item in the resilience measurement table, in response to a configuration operation on the chaos engineering platform, configure a fault scenario, a stress scenario, and a data item to be monitored corresponding to the test item in the resilience measurement of the target distributed system;

[0038] A fault drill module, configured to perform a fault drill on the target distributed system according to the fault scenario and the stress scenario, and collect test data corresponding to the data item to be monitored;

[0039] A determination module, configured to determine a resilience test result of the target distributed system under the test item according to the test data;

[0040] A display module, configured to display the resilience test result under the test item at the corresponding display position of the resilience measurement table.

[0041] An embodiment of the present application further provides an electronic device, including: a processor, a memory, and a bus, where the memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of a resilience measurement method for a distributed system based on chaos engineering as described above are executed.

[0042] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the steps of a resilience measurement method for a distributed system based on chaos engineering as described above are executed.

[0043] A method and device for measuring the resilience of a distributed system based on chaos engineering provided by an embodiment of the present application can intuitively present the results of the resilience test of the distributed system through a resilience measurement table displayed in a user interface provided by a chaos engineering platform; for each test item in the resilience measurement table, the chaos engineering platform performs a fault drill of chaos engineering by configuring a fault scenario and a stress scenario, and can simulate various faults to conduct a comprehensive fault drill on the distributed system.

[0044] Among them, the simulation of faults does not require modifying the program code, manually inputting computer instructions or running corresponding program scripts, and does not require manual operations such as physically starting and stopping the server hardware or cutting off the network cable. The entire implementation process only requires simple configuration operations in the chaos engineering platform to execute the chaos engineering fault drill and achieve various fault injections, thereby improving the integrity and efficiency of the resilience measurement of the distributed system.

[0045] On the other hand, according to the user's selection operation, a resilience measurement table suitable for different distributed systems can be constructed as needed, and the resilience measurement progress and weak test index dimensions are displayed in real time in the resilience measurement table, enabling the user to intuitively perceive the resilience measurement situation of the distributed system.

[0046] To make the above objects, features, and advantages of the present application more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, details are described as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0048] Figure 1 Shows a schematic diagram of a resilience measurement table of a distributed system provided by an embodiment of the present application;

[0049] Figure 2 Shows one of the flowcharts of a method for measuring the resilience of a distributed system based on chaos engineering provided by an embodiment of the present application;

[0050] Figure 3 Shows another flowchart of a method for measuring the resilience of a distributed system based on chaos engineering provided by an embodiment of the present application;

[0051] Figure 4 Shows a schematic structural diagram of a device for measuring the resilience of a distributed system based on chaos engineering provided by an embodiment of the present application;

[0052] Figure 5 The structural schematic diagram of an electronic device provided by an embodiment of the present application is shown. Detailed implementation manners

[0053] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part rather than all of the embodiments of the present application. Components of the embodiments of the present application generally described and illustrated in the accompanying drawings herein can be arranged and designed in a variety of different configurations. Therefore, the detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application claimed, but merely represents selected embodiments of the present application. Based on the embodiments of the present application, every other embodiment obtained by those skilled in the art without making creative efforts belongs to the scope of protection of the present application.

[0054] Resilience has different cognitive forms at different stages of the development of social productive forces. Originally, it referred to the requirement for the resilience of metal materials and had a relatively strong simple linear correlation with engineering mechanics. Assuming that the system has only one stable state, and resilience is the ability of the system to return to this stable state after being impacted, which is called engineering resilience. With the development of information technology, resilience (also translated as elasticity and recovery force) in the embodiments of the present application refers to the ability of a system, organization or individual to adapt to, resist and recover in the face of uncertainty, impact and pressure, and refers to the ability of an information system to continuously operate, continuously resist threats and impacts, and maintain growth in a changing environment. Therefore, for a distributed system, resilience measurement is of great significance.

[0055] However, it has been found through research that the prior art generally indirectly measures resilience by simulating faults through manual command input or executing program scripts to test the system to examine the availability or reliability of the distributed system. However, this test method relying on manual operation has low efficiency and a high error probability; it can only simulate very limited types of faults, resulting in an incomplete measurement of the resilience of the distributed system, and the measurement results are difficult to intuitively perceive.

[0056] Based on this, the embodiments of the present application provide a method for measuring the resilience of a distributed system based on chaos engineering to intuitively present the results of the resilience test of the distributed system and improve the integrity and efficiency of the resilience measurement of the distributed system.

[0057] The resilience measurement method provided by the embodiments of the present application is applied to a chaos engineering platform. The chaos engineering platform applies chaos engineering-related technologies, can simulate various system failures of a distributed system, and comprehensively measures the resilience of the distributed system by actively injecting faults. Such as: service start / stop, downtime, network card downtime, network latency, network packet loss, application stop, service stop, process suspension, JVM return value tampering, disk I / O exception, Druid database connection pool full, data center failure cluster switchover, CPU full load, power module failure, container crash, microservice call latency, message queue exception, memory overflow, etc.

[0058] Among them, the chaos engineering platform can run on various terminal devices to provide a visual user interaction interface through the terminal device to perform data interaction with users, such as receiving various configuration operations of users and real-time displaying information related to the resilience measurement of the distributed system. Specifically, the user interaction interface at least displays a resilience measurement table of the target distributed system. The target distributed system is the distributed system to be measured for resilience. In the embodiments of the present application, the resilience measurement table visually shows the test items of the target distributed system in a visual manner; and as the fault drill process progresses, fills in the test results under different test items of the target distributed system, enabling users to intuitively perceive the currently completed test items and test results and the unexecuted test items, so as to understand the resilience measurement process.

[0059] Please refer to Figure 1 , Figure 1 which is a schematic diagram of a resilience measurement table of a distributed system provided by the embodiments of the present application. In one example, for a target distributed system with virtual machines as the service cluster, the resilience measurement table is as shown in Figure 1 . Among them, the test index dimensions include: test dimensions of high availability such as cluster high availability, link high availability, consistency, batch service high availability, database cluster effectiveness, message service high availability, failover effectiveness, service registration effectiveness, etc.; under the test index dimension of cluster high availability, there are index sub-dimensions of cluster ability, load balancing ability, and scaling ability; under the index sub-dimension of cluster ability, there are test items of whether it is a cluster, one-key start effectiveness, self-pulling ability, full node failure recovery ability, partial node failure recovery ability, node failure detection and warning ability, and threshold effectiveness.

[0060] Before the resilience measurement begins, all test items are initialized to the "not yet tested" state, and the "not yet tested" state is filled in the display position corresponding to each test item in the resilience measurement form. For example, the "○" symbol can be used to fill in the test result column corresponding to each test item in the resilience measurement form to indicate that the test item has not been tested; as the test progresses, for each tested test item, update the test result column according to the corresponding test result, so as to visually display the resilience measurement result.

[0061] In a possible implementation manner, the resilience measurement form of the target distributed system can be constructed in the following way:

[0062] In response to an import operation for the chaos engineering platform, import a pre-constructed original resilience measurement form into the chaos engineering platform and display it on the user interface; in response to a selection operation for the original resilience measurement form in the chaos engineering platform, select test index dimensions, index sub-dimensions under each test index dimension, and test items under each index sub-dimension from the original resilience measurement form to construct the resilience measurement form of the target distributed system.

[0063] Here, through extensive analysis of various distributed systems, developers can comprehensively determine the test index dimensions involved in the resilience measurement of the distributed system, the index sub-dimensions under each test index dimension, and the test items under each index sub-dimension, and then construct and import the original resilience measurement form. After that, through the import operation, developers can import the original resilience measurement form into the chaos engineering platform and display it on the user interface.

[0064] When performing resilience measurement on a specific actual target distributed system, due to the different system architectures of different distributed systems and the different business requirements of different customers, in response to a selection operation for the original resilience measurement form, test index dimensions, index sub-dimensions under each test index dimension, and test items under each index sub-dimension can be selected from the original resilience measurement form to construct the resilience measurement form of the target distributed system from the original resilience measurement form. In this way, it can meet the resilience measurement of distributed systems under different system architectures and different business requirements of users.

[0065] Please refer to Figure 2 , Figure 2 which is one of the flowcharts of a method for measuring the resilience of a distributed system based on chaos engineering provided by an embodiment of the present application. As Figure 2 shown in

[0066] S101. For each test item in the resilience measurement form, in response to a configuration operation on the chaos engineering platform, configure the fault scenario, stress scenario, and data items to be monitored corresponding to this test item in the resilience measurement of the target distributed system.

[0067] In this step, the chaos engineering platform can automatically select untested test items from the resilience measurement form in sequence to conduct fault drill tests on the target distributed system, or in response to the user's selection operation, select test items from the resilience measurement form. After selecting a test item, a corresponding configuration interface for the test item is displayed in the user interface. Under different test items, the user can, through simple configuration operations, configure different fault scenarios, stress scenarios, and data items to be monitored to test different system capabilities of the target distributed system, avoiding the problems of low efficiency and error-proneness caused by manually writing code instructions.

[0068] In the embodiment of the present application, the chaos engineering platform creates faults by actively introducing abnormal states of software or hardware into the system, and determines the resilience measurement result based on the behavior performance of the system under various pressures. Among them, the fault scenario is used to set the method for injecting faults into the distributed system during the execution of the drill, so as to create faults and inject them into the system for drill according to the fault scenario; the stress scenario is used to set the method for applying stress to the distributed system, so as to apply the expected experimental stress to the distributed system according to the stress scenario; the data items to be monitored refer to the data content that needs to be monitored to measure the system resilience, so as to determine the resilience measurement result according to the data values of the data items to be monitored during the experiment.

[0069] In specific implementation, step S101 may include:

[0070] S1011. For each test item in the resilience measurement form, in response to a first configuration operation on the chaos tool in the chaos engineering platform, select the types of fault events under this test item, and configure the fault drill objects and fault parameters corresponding to each selected fault event in the target distributed system, and generate a fault scenario corresponding to each fault event.

[0071] Here, there is a bottom-layer chaos tool in the chaos engineering platform, and the user can configure the running parameters of the chaos tool through operations on the interaction interface. There is at least one selectable fault event under each test item. In response to the user's first configuration operation, at least one fault event can be selected from under the test item, and the fault drill objects and fault parameters of each selected fault event can be configured to generate a fault scenario corresponding to each fault event.

[0072] Among them, the objects of fault drill can be physical hosts, virtual machines, container nodes / pods / containers, physical machine clusters, virtual machine clusters, and container clusters in the target distributed system, etc.; the embodiments of the present application also provide more fine-grained fault drills, such as service instances running on a certain node in the cluster. Optional types of fault events include, but are not limited to: killing processes, pausing processes, system downtime and restart, network packet loss, network latency, network storage exception, high CPU occupancy, memory shortage, and local storage exception. The parameters of the fault event include: the start time of the fault drill, the end time, and the running time of the event, etc. According to the fault drill object, each fault event and the corresponding parameters, a fault scenario is generated and saved.

[0073] Corresponding to the above example, for the test item of "cluster ability - self-pulling ability" in the above-mentioned resilience measurement table. When creating a fault scenario, the optional types of fault events include: forced process offline and process suspension. When the user selects the fault event of "forced process offline", first, fill in the name, duration, and description information of the fault drill; second, select a certain service instance in the cluster as the object of the fault drill, and the IP of the node where the service instance is located and the process number of the service instance need to be specified in the chaos engineering platform; finally, specify the parameters of the fault event, such as the running time. In this way, the fault scenario corresponding to the fault event of "forced process offline" is generated and saved, waiting to be executed.

[0074] S1012. In response to a second configuration operation for the test tool in the chaos engineering platform, configure the pressure parameters corresponding to each fault event, and generate a pressure application scenario corresponding to each fault event according to the pressure parameters and the fault drill object.

[0075] Here, test tools such as LoadRunner, JMeter, and APTS can also be integrated in the chaos engineering platform. The user can configure the running parameters of the test tool through operations on the interaction interface. Specifically, for each fault event, the user can configure the load pressure parameters according to the system processing capacity of the target distributed system. For example, 50% of the maximum processing capacity of the target distributed system is used as the load pressure parameter, and a corresponding pressure application scenario is generated according to the specified fault drill object. Then save it and wait to be executed.

[0076] S1013. In response to a third configuration operation for the system monitoring platform in the chaos engineering platform, determine the resource data and business metrics related to each fault event as the data items to be monitored; and in response to a fourth configuration operation for the service registration and discovery component in the chaos engineering platform, determine the service instance status parameters of the fault drill object as the data items to be monitored.

[0077] Here, during the fault drill, it is necessary to observe and record the system data and comprehensive performance. The chaos engineering platform can also integrate system monitoring platforms such as APM, ZABBIX, Prometheus, Grafana, etc., as well as service registration and discovery components such as Nacos, Eureka, and ZooKeeper to comprehensively monitor the target distributed system from multiple perspectives. Through operations on the interactive interface, users can configure the operating parameters of the system monitoring platform and service registration and discovery components, that is, select the data items to be monitored. Here, in the embodiments of the present application, the resource data, business metrics, and service instance status parameters of the fault drill object related to each fault event are selected as the data items to be monitored. Exemplarily, the resource data may include basic resource conditions such as the CPU utilization rate, memory occupancy rate, and IO busyness of each host; the business metrics may include the transaction failure rate, response time, number of transactions per second, etc.; the service instance status parameters may include: fault injection time, recovery time of the number of transactions per second, actual offline time of the service instance, service instance recovery time, etc.

[0078] Optionally, the above-mentioned test tool, system monitoring platform, and service registration and discovery components are all integrated in the chaos engineering platform, which is convenient for management and enables the chaos engineering platform to have a more comprehensive fault drill function; or, the test tool, system monitoring platform, and service registration and discovery components can also be set independently of the chaos engineering platform to meet personalized business needs; at this time, through the data connection between the chaos engineering platform and these components, users can also perform corresponding configurations through configuration operations on the interactive interface.

[0079] S102. Perform a fault drill on the target distributed system according to the fault scenario and the stress scenario, and collect the test data corresponding to the data items to be monitored.

[0080] In this step, if the fault parameters of the fault scenario include an execution time and the stress parameters of the stress scenario include an execution time; then at the corresponding time, the chaos engineering platform can automatically perform a fault drill according to the fault scenario and the stress scenario. Alternatively, in response to a start instruction manually triggered by the user, the chaos engineering platform performs a fault drill according to the fault scenario and the stress scenario. In specific implementation, the chaos engineering platform applies pressure to the target distributed system according to the stress parameters set in the stress scenario, and under the pressure background (such as after the stress scenario runs for a certain period of time), injects fault events into the target distributed system according to the fault parameters of the fault scenario to perform a fault drill; and continuously collects the test data of the target distributed system according to the configured data items to be monitored.

[0081] In specific implementation, step S102 may include:

[0082] S1021. For each fault event under the test project, the test tool sends a data stream to the fault drill object according to the pressure parameters corresponding to the pressure application scenario for pressure application.

[0083] Exemplarily, using JMeter as the pressure - sending engine, taking 50% of the maximum processing capacity of the target distributed system as the load pressure, a large number of request links are sent to a service instance on a certain node in the target distributed system for pressure application, and the target distributed system is waited to run stably for a certain period of time, such as 5 minutes, under the pressure application scenario.

[0084] S1022. During the pressure application process, the chaos tool sends a fault injection instruction to the fault drill object according to the fault parameters corresponding to the fault scenario, so that the Agent probe pre - installed in the target distributed system executes fault event injection on the fault drill object according to the fault injection instruction.

[0085] Corresponding to the above example, for "cluster ability - self - pulling - up ability" in the above - mentioned resilience measurement table, a chaos engineering fault drill is executed. The chaos tool sends a fault injection instruction to the specified service instance according to the fault parameters corresponding to the fault scenario to inject a "process forced offline" fault event. After receiving the fault injection instruction, the Agent probe pre - installed in the target distributed system simulates the corresponding fault according to the fault injection instruction, executes fault event injection on the specified service instance, and can control the continuous operation of the fault event for a certain period of time.

[0086] S1033. The system monitoring platform and the service registration and discovery component collect the test data corresponding to the data items to be monitored of the fault drill object.

[0087] Corresponding to the above example, for "cluster ability - self - pulling - up ability" in the above - mentioned resilience measurement table, after the fault event injection, the system monitoring platform observes whether the transaction response time and the number of transactions per second in the monitored target distributed system rise and fall as expected. The service registration and discovery component monitors the change of the registration status of the service instance, and records the actual offline time T1 of the service instance and the actual recovery time T2 of the service instance.

[0088] S103. Determine the resilience test result of the target distributed system under this test project according to the test data.

[0089] In this step, by analyzing the target distributed system and according to the system architecture and business requirements, the resilience measurement table in the embodiments of the present application pre-configures steady-state assumptions for each type of failure event under each test item; the steady-state assumption refers to the specific index threshold or signal basis for determining whether each test item of the system passes during the test, and is used as the test verification standard for subsequent failure drills through chaos engineering. By verifying the test data according to the steady-state assumption, the resilience test results of the target distributed system under each test item can be determined.

[0090] Exemplarily, for whether the test of the "cluster ability - self-pulling ability" test item in the above-mentioned resilience measurement table passes or not, it needs to be determined according to the length of the self-pulling time of a certain service instance in the application cluster of the test distributed system after being injected with failure events corresponding to different failure scenarios. Therefore, for the failure event of "forced process offline", the corresponding steady-state assumption can be set as "the service instance will self-pull within 60 seconds after being forced offline", that is, "the self-pulling time threshold is 60 seconds".

[0091] Therefore, optionally, in addition to the above implementation manner of determining the data items to be monitored by the user's configuration operation, the chaos process platform can also automatically parse the corresponding data items to be monitored according to the steady-state assumptions corresponding to each test item in the resilience measurement table, and automatically configure the system monitoring platform and the service registration and discovery component. In this way, the automation degree and implementation efficiency of the resilience measurement of the distributed system are further improved. For example, by parsing "the service instance will self-pull within 60 seconds after being forced offline" through natural language related processing technology, it is determined that the parameter constrained by the steady-state assumption is the self-pulling time, and then the data item to be monitored corresponding to the parameter constrained by the steady-state assumption can be determined according to the preset calculation formula in the platform, such as the self-pulling time corresponding to the actual offline time of the service instance and the actual recovery time of the service instance. In this way, when the parameters constrained by the steady-state assumptions of some test items change, only the resilience measurement table needs to be modified, and the data items to be monitored can be automatically changed accordingly, without the user having to re-configure.

[0092] Furthermore, when there are multiple failure events under any test item, step S103 may include:

[0093] S1031. For each type of failure event under this test item, according to the test data and the steady-state assumption configured for this type of failure event in the resilience measurement table, determine the sub-resilience test result of the target distributed system under this type of failure event; S1032. Synthesize the sub-resilience test results under each type of failure event to determine the resilience test result of the target distributed system under this test item.

[0094] In a possible implementation, when the steady-state hypothesis is the test verification standard for specific index threshold categories, step S1031 may include:

[0095] For each failure event under this test item, analyze the steady-state hypothesis configured for this failure event, and parse out the target parameters constrained by the steady-state hypothesis and the numerical judgment conditions of the target parameters; determine the target monitored data item corresponding to the target parameter and the calculation relationship between the target monitored data items; extract the target test data corresponding to the target monitored data item from the test data, and obtain the parameter value of the target parameter according to the calculation relationship; according to the parameter value of the target parameter and the numerical judgment conditions of the target parameter, determine the sub-resilience test result of the target distributed system under this failure event.

[0096] Exemplarily, for the test item "Cluster Capacity - Self-Pulling-Up Capacity" in the above-mentioned resilience metric table, for one of the failure events, "Process Forced Offline", the preset steady-state hypothesis condition is that "the service instance will self-pull up within 60 seconds after being forced offline"; through natural language processing, regular expressions, dictionary library matching, etc., it can be analyzed that the target parameter constrained by the steady-state hypothesis is: the actual self-pulling-up time △t of the service instance; the numerical judgment condition of the target parameter: less than 60 seconds. According to the pre-stored calculation formula library in the chaos engineering platform, it can be determined that the target monitored data item corresponding to the target parameter "the actual self-pulling-up time △t of the service instance" is: the actual offline time T1 of the service instance and the actual recovery time T2 of the service instance, and the calculation relationship between the target monitored data items is subtraction, that is, △t = T2 - T1. Then, extract the target test data corresponding to the target monitored data item from the test data, that is, the values of T1 and T2, and obtain the parameter value of the target parameter △t according to the calculation relationship. Finally, according to the parameter value of the target parameter △t and the numerical judgment condition of the target parameter: less than 60 seconds, determine the sub-resilience test result of the target distributed system under the "Process Forced Offline" failure event. The sub-resilience test result includes pass and fail.

[0097] Specifically, if △t < 60 seconds, the steady-state hypothesis holds, and the sub-resilience test result under this failure event is pass, indicating that the cluster has an effective self-pulling-up ability in the "Process Forced Offline" failure scenario. On the contrary, if △t > 60 seconds, the steady-state hypothesis does not hold, and the sub-resilience test result under this failure event is fail, indicating that the cluster does not have an effective self-pulling-up ability in the "Process Forced Offline" failure scenario.

[0098] In this way, the chaos engineering platform can automatically obtain the resilience test results of the test project based on the monitored test data and the pre-configured resilience metric table, so as to improve the automation degree and efficiency of the resilience measurement of the distributed system.

[0099] For another type of failure event, "process suspension", continue to conduct a drill on the "process suspension" failure event for a certain service instance of the cluster, and determine the corresponding sub-resilience test results according to the corresponding steady-state assumptions. Finally, based on the sub-resilience test results under the two types of failure events, determine the resilience test results of the target distributed system under this test project. Exemplarily, the resilience test results can be all passed, all failed, and partially passed. That is, when the sub-resilience test results under all failure events of a test project are all passed, the resilience test result of the test project is all passed; when the sub-resilience test results under all failure events of a test project are all not passed, the resilience test result of the test project is all failed; when the sub-resilience test results under some failure events of a test project are all passed, the resilience test result of the test project is partially passed. Further, the resilience test results of the test project can also be comprehensively obtained according to the combination of logical operators such as "AND", "OR", and "NOT" based on the sub-resilience test results under each type of failure event.

[0100] S104. Display the resilience test results under this test project at the corresponding display position in the resilience metric table.

[0101] In this step, update the resilience metric table according to the resilience test results under this test project, that is, fill the resilience test results into the corresponding display positions under the test project.

[0102] Corresponding to the above example, based on the sub-resilience test results of the two failure drills of "forced process offline" and "process suspension", update the mark under the "cluster ability - self-pulling ability" test project in the resilience metric table. For example, update the original "○" symbol indicating not yet tested to: all passed: √, all failed: ×, partially passed: In this way, users can intuitively perceive the currently completed test projects and test results.

[0103] Further, when there are multiple types of failure events under any test project, the corresponding resilience test results of the test project include passed, failed, and partially passed. The method further includes:

[0104] Mark the test item as a completed test item; determine the resilience measurement progress of the target distributed system according to the proportion of the completed test items in the resilience measurement table, and display the resilience measurement progress at the corresponding display position of the resilience measurement progress item in the resilience measurement table; determine the weak test index dimension of the target distributed system according to the test items with unqualified and partially qualified resilience measurement results in the resilience measurement table.

[0105] In this way, the "resilience measurement table" is used as a tool to present the work progress of the resilience verification of the distributed system. According to the number and proportion of the test items marked with the "completed test" identifier in the current resilience measurement table, the current progress of the resilience verification work is intuitively presented. On the other hand, according to the number and proportion of the test items marked with "all passed" in the resilience measurement table, the resilience status of the system can be intuitively presented, and according to the test items marked with "all failed" and "partially passed" in the table, the weak test index dimension of the system can be intuitively understood, so that developers can carry out targeted optimization.

[0106] Further, after displaying the resilience test result under the test item at the corresponding display position of the test item in the resilience measurement table, the following is further included:

[0107] Cancel the fault drill executed on the target distributed system; determine the system recovery situation of the target distributed system according to the system data corresponding to the data item to be monitored; after the system recovery situation indicates that the target distributed system has returned to normal, execute the resilience measurement method for the next test item.

[0108] Here, after the fault drill for a certain test item is completed, the chaos tool of the chaos engineering platform releases the fault injection, restores the faulty node, and the scenario runs continuously for 5 minutes to observe the recovery situation of each business indicator in a short time. The goal is to restore the normal operation of the target distributed system service and prepare for the fault drill of other test items later. After determining that the target distributed system has returned to normal according to the system recovery situation, the resilience measurement method for the next test item can be executed.

[0109] In this way, by constructing a resilience measurement table, the test results and test progress can be intuitively presented; for each test item in the resilience measurement table, by configuring the fault scenario and stress scenario of the fault drill process to execute the fault drill of the chaos engineering, the integrity and efficiency of measuring the resilience of the distributed system can be improved.

[0110] Please refer to Figure 3 , Figure 3 which is the second flowchart of a method for measuring the resilience of a distributed system based on chaos engineering provided by another embodiment of this application. As Figure 3As shown, the resilience measurement method provided by the embodiments of the present application is applied to a chaos engineering platform. The chaos engineering platform uses a system resilience measurement table as a presentation tool. First, a resilience measurement table for the target distributed system is created. In the table, various resilience test index dimensions are listed, and the steady-state assumptions corresponding to each test item are set. The chaos engineering platform conducts system resilience measurement tests for each test item in the manner of chaos engineering fault drills.

[0111] Specifically, automatically select an untested test item in the resilience measurement table, and determine each fault event involved in the test item; create a fault scenario and a stress scenario corresponding to each fault event; preset the data items to be monitored by the system monitoring platform and the service registration and discovery component; execute the chaos engineering fault scenario under the background pressure of the stress scenario, and observe and record the data items to be monitored of the key system indicators and data; after that, when the drill ends, restore the system; compare the test data with the preset steady-state assumptions to obtain the binary sub-resilience test results of whether the test passes under each fault event; determine whether all fault events have been drilled under this test item; if not, return to execute the fault drill for the next fault event; if so, comprehensively determine the resilience test results under this test item based on the passing situations of the sub-resilience test results under various fault events, and fill the resilience test results into the system resilience measurement table for intuitive presentation.

[0112] After that, determine whether there are untested test items in the resilience measurement table; if so, return to conduct system resilience measurement tests for the untested test items; if not, end a distributed system resilience measurement, and then a resilience analysis report can be generated based on the filled resilience measurement table to comprehensively analyze the resilience situation of the distributed system.

[0113] The resilience measurement method of the distributed system based on chaos engineering provided by the embodiments of the present application has the following beneficial effects: at the level of the presentation tool for system resilience measurement, through the resilience measurement table displayed in the user interface provided by the chaos engineering platform, the resilience status of the system and the weak links of the system can be presented intuitively and completely; at the level of controlling the test progress of system resilience measurement, through the resilience measurement table, relevant staff can clearly perceive the progress of the system test work at a glance; in terms of the test during system resilience measurement, it is achieved through chaos engineering fault drills. The simulation of the fault scenario does not require much professional cooperation, does not require modifying the program code, does not require manual input of computer instructions or running corresponding program scripts, and does not require physical start / stop of the server hardware, cutting off the network cable, etc. The entire implementation process only requires selecting the fault drill object, the type and parameters of the fault to be injected in the chaos engineering platform, and then starting the chaos engineering fault drill with one key to achieve various fault injections and verify the test items.

[0114] Please refer toFigure 4 , Figure 4 is a schematic structural diagram of a resilience measurement device for a distributed system based on chaos engineering provided by an embodiment of the present application. As Figure 4 shown in

[0115] The configuration module 410 is configured to, for each test item in the resilience measurement table, in response to a configuration operation for the chaos engineering platform, configure a failure scenario, a stress scenario, and a data item to be monitored corresponding to the test item in the resilience measurement of the target distributed system;

[0116] The failure drill module 420 is configured to perform a failure drill on the target distributed system according to the failure scenario and the stress scenario, and collect test data corresponding to the data item to be monitored;

[0117] The determination module 430 is configured to determine a resilience test result of the target distributed system under the test item according to the test data;

[0118] The display module 440 is configured to display the resilience test result under the test item at a display position corresponding to the test item in the resilience measurement table.

[0119] It should be noted that the specific implementation manner of the device can be referred to the method embodiment. Since the principle of the resilience measurement device in the embodiment of the present application for solving problems is similar to the above-mentioned resilience measurement method in the embodiment of the present application, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be elaborated.

[0120] Please refer to Figure 5 , Figure 5 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 5 shown in

[0121] The memory 520 stores machine-readable instructions executable by the processor 510. When the electronic device 500 runs, the processor 510 communicates with the memory 520 through the bus 530. When the machine-readable instructions are executed by the processor 510, the steps of a resilience measurement method for a distributed system based on chaos engineering in the method embodiment as shown above can be executed. The specific implementation manner can be referred to the method embodiment and will not be elaborated here. Figure 2 An embodiment of the present application further provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, the steps of a resilience measurement method for a distributed system based on chaos engineering in the method embodiment as shown above can be executed.

[0122] When the computer program is run by a processor, the steps of a resilience measurement method for a distributed system based on chaos engineering in the method embodiment as shown above can be executed. The specific implementation manner can be referred to the method embodiment and will not be elaborated here. Figure 2The steps of a resilience measurement method for a distributed system based on chaos engineering in the method embodiments shown can be specifically implemented by referring to the method embodiments and will not be elaborated here.

[0123] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.

[0124] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. Also, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0125] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0126] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0127] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0128] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present application, which are used to illustrate the technical solutions of the present application, rather than limiting it. The protection scope of the present application is not limited thereto. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present application can still modify the technical solutions recorded in the foregoing embodiments or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A resilience measurement method for a distributed system based on chaos engineering, characterized in that, Applied to a chaos engineering platform, the chaos engineering platform provides a user interaction interface; The user interaction interface displays a resilience measurement table of a target distributed system; The method includes: For each test item in the resilience measurement table, in response to a configuration operation on the chaos engineering platform, configure the fault scenario, stress scenario, and monitored data item corresponding to this test item in the resilience measurement of the target distributed system; According to the fault scenario and the stress scenario, perform a fault drill on the target distributed system and collect the test data corresponding to the monitored data item; Determine the resilience test result of the target distributed system under this test item according to the test data; Display the resilience test result under this test item at the display position corresponding to this test item in the resilience measurement table; The method further includes: Automatically parse the parameters constrained by the steady-state hypothesis according to the steady-state hypothesis corresponding to each test item in the resilience measurement table; and according to the pre-stored calculation formula library of the chaos engineering platform, determine the monitored data item corresponding to the parameters constrained by the steady-state hypothesis; The resilience measurement table is pre-configured with the steady-state hypothesis for each type of fault event under each test item; The steady-state hypothesis refers to the specific index threshold or signal basis for determining whether each test item of the system passes during the test; Automatically configure the system monitoring platform and the service registration and discovery component so that the system monitoring platform and the service registration and discovery component collect the test data corresponding to the monitored data item.

2. The method according to claim 1, wherein For each test item in the resilience measurement table, in response to a configuration operation on the chaos engineering platform, configure the fault scenario, stress scenario, and monitored data item corresponding to this test item in the resilience measurement of the target distributed system, including: For each test item in the resilience measurement table, in response to a first configuration operation on the chaos tool in the chaos engineering platform, select the type of fault event under this test item, and configure the fault drill object and fault parameters corresponding to each selected fault event in the target distributed system to generate a fault scenario corresponding to each fault event; In response to a second configuration operation on the test tool in the chaos engineering platform, configure the pressure parameter corresponding to each fault event, and generate a stress scenario corresponding to each fault event according to the pressure parameter and the fault drill object; In response to a third configuration operation on the system monitoring platform in the chaos engineering platform, determine the resource data and business metrics related to each fault event as the monitored data item; and in response to a fourth configuration operation on the service registration and discovery component in the chaos engineering platform, determine the service instance status parameter of the fault drill object as the monitored data item.

3. The method according to claim 2, characterized in that, According to the fault scenario and the stress scenario, perform a fault drill on the target distributed system and collect the test data corresponding to the monitored data item, including: For each fault event under this test item, the test tool sends a data stream to the fault drill object for stress according to the pressure parameter corresponding to the stress scenario; During the pressure application process, the chaos tool sends a fault injection instruction to the fault exercise object according to the fault parameters corresponding to the fault scenario, so that the Agent probe pre-installed in the target distributed system executes a fault event injection on the fault exercise object according to the fault injection instruction; The system monitoring platform and the service registration and discovery component collect the test data corresponding to the to-be-monitored data items of the fault exercise object.

4. The method according to claim 1, wherein Determine the resilience test result of the target distributed system under this test item according to the test data, including: For each fault event under this test item, determine the sub-resilience test result of the target distributed system under this fault event according to the test data and the steady-state assumptions configured for this fault event in the resilience measurement table; Integrate the sub-resilience test results under each fault event to determine the resilience test result of the target distributed system under this test item.

5. The method according to claim 4, wherein For each fault event under this test item, determine the sub-resilience test result of the target distributed system under this fault event according to the test data and the steady-state assumptions configured for this fault event in the resilience measurement table, including: For each fault event under this test item, analyze the steady-state assumptions configured for this fault event, and parse out the target parameters constrained by the steady-state assumptions and the numerical judgment conditions of the target parameters; Determine the target to-be-monitored data item corresponding to the target parameter and the calculation relationship between the target to-be-monitored data items; Extract the target test data corresponding to the target to-be-monitored data item from the test data, and obtain the parameter value of the target parameter according to the calculation relationship; Determine the sub-resilience test result of the target distributed system under this fault event according to the parameter value of the target parameter and the numerical judgment conditions of the target parameter.

6. The method according to claim 1, wherein Construct the resilience measurement table of the target distributed system in the following way: In response to the import operation for the chaos engineering platform, import the pre-constructed original resilience measurement table into the chaos engineering platform and display it on the user interface; In response to the selection operation for the original resilience measurement table in the chaos engineering platform, select the test index dimension, the index sub-dimension under each test index dimension, and the test item under each index sub-dimension from the original resilience measurement table to construct the resilience measurement table of the target distributed system.

7. The method according to claim 6, wherein The resilience test result includes all passed, all not passed, and partially passed; the method further includes: Mark this test item as a completed test item; Determine the resilience measurement progress of the target distributed system according to the proportion of the completed test items in the resilience measurement table, and display the resilience measurement progress at the display position corresponding to the resilience measurement progress item in the resilience measurement table; Determine the weak test index dimension of the target distributed system according to the test items in the resilience measurement table with the resilience measurement results of all not passed and partially passed.

8. The method according to claim 1, wherein After displaying the resilience test results under the test item at the corresponding display position in the resilience measurement table, the method further includes: lifting the fault drill performed on the target distributed system; determining the system recovery status of the target distributed system according to the system data corresponding to the data item to be monitored; after the system recovery status indicates that the target distributed system has returned to normal, performing the resilience measurement method for the next test item.

9. A resilience measurement device for a distributed system based on chaos engineering, characterized in that, Applied to a chaos engineering platform, the chaos engineering platform provides a user interaction interface; the user interaction interface displays a resilience measurement table of the target distributed system; the device includes: a configuration module, configured to, for each test item in the resilience measurement table, in response to a configuration operation on the chaos engineering platform, configure the fault scenario, the stress scenario, and the data item to be monitored corresponding to the test item in the resilience measurement of the target distributed system; a fault drill module, configured to perform a fault drill on the target distributed system according to the fault scenario and the stress scenario, and collect test data corresponding to the data item to be monitored; a determination module, configured to determine the resilience test result of the target distributed system under the test item according to the test data; a display module, configured to display the resilience test result under the test item at the corresponding display position in the resilience measurement table; the configuration module is further configured to: automatically parse out the parameters constrained by the steady-state hypothesis according to the steady-state hypothesis corresponding to each test item in the resilience measurement table; and determine the data item to be monitored corresponding to the parameters constrained by the steady-state hypothesis according to the calculation formula library pre-stored in the chaos engineering platform; the steady-state hypothesis is pre-configured for each fault event under each test item in the resilience measurement table; the steady-state hypothesis refers to the specific index threshold or signal basis for determining whether each test item of the system passes during the test; automatically configure the system monitoring platform and the service registration and discovery component, so that the system monitoring platform and the service registration and discovery component collect the test data corresponding to the data item to be monitored.

10. An electronic device, characterized in that, including: a processor, a memory, and a bus, the memory stores machine-readable instructions executable by the processor, when the electronic device runs, the processor communicates with the memory through the bus, and when the machine-readable instructions are run by the processor, the steps of a method for measuring the resilience of a distributed system based on chaos engineering as described in any one of claims 1 to 8 are executed.

Citation Information

Patent Citations

  • Fault injection method and device, storage medium and terminal

    CN116909787A

  • System load balancing capability verification method and device based on chaos engineering

    CN117056081A