Chaotic engineering fault arrangement and injection method and device, electronic equipment and medium
By obtaining system information and calling fault type prediction models and libraries, and automatically orchestrating and injecting faults, the problem of incomplete coverage of fault scenarios in the existing technology is solved, and the efficiency of system failure self-healing evaluation is improved.
Patent Information
- Application Number
- CN202510412435.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-18
AI Technical Summary
When the existing technology injects faults, it cannot effectively cover all kinds of fault scenarios, and it depends on the experience of operation and maintenance personnel, resulting in low response verification efficiency of the application system.
By obtaining the system name information, fault range information and expected recovery time entered by the user, the fault type prediction model is called and combined with the pre-generated fault library and expert library, the fault is automatically arranged and injected to perform system fault self-healing assessment.
It has achieved the reduction of manual intervention on the basis of covering all kinds of fault scenarios as much as possible, and improved the degree of automation of fault injection and the efficiency of system fault self-healing evaluation.
Smart Images

Figure CN120342844A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of fault injection, and particularly to a method, device, electronic device and medium for chaos engineering fault orchestration and injection. Background Art
[0002] With the development of business, the architecture of the system will also be updated and iterated accordingly. Especially in the distributed system architecture, the number of services increases, the call chain becomes longer, and the dependency relationships between services become increasingly complex. It is difficult to evaluate the impact of a single service failure on the entire system. During the high-speed iteration of the system, how to continuously ensure the stability of the system has always been a great challenge.
[0003] The idea of chaos engineering is to actively inject faults to discover abnormal points in the system in advance. Once the abnormal points are discovered, they can be improved, thereby continuously enhancing the stability of the system. Implementing chaos engineering can discover problems in the production environment early, improve the efficiency of fault emergency handling, sort out the strong and weak dependency relationships between services, verify the effectiveness of the emergency plan, etc., and gradually build a highly available system with resilience. Therefore, how to reasonably inject faults into the system is particularly important.
[0004] At present, the method of fault injection usually injects faults into application services, without considering the scenarios where basic resources such as memory, network, and CPU may fail, or both the target scenario type and the target fault scenario are preset, and users choose according to actual operation and maintenance experience, and the covered scenarios are not comprehensive enough. The fault parameters and experimental operation parameters are configured by users in advance, and the automatic configuration of fault scenarios and parameters cannot be realized. The design and coverage of scenarios rely heavily on the experience of operation and maintenance personnel, and the reaction verification of the application system has poor targetability, resulting in low reaction verification efficiency of the application system. Summary of the Invention
[0005] The technical problem to be solved by the embodiments of this application is to provide a method, device, electronic device and medium for chaos engineering fault orchestration and injection, so as to effectively automate the orchestration and injection of faults by identifying the intention of fault injection input by the user, and reduce the intervention of manual scenario orchestration on the basis of covering various fault scenarios as much as possible.
[0006] In the first aspect, the embodiments of this application provide a method for chaos engineering fault orchestration and injection, and the method includes:
[0007] Obtain the system name information, fault scope information and expected recovery time input by the user;
[0008] Call the fault type prediction model to process the system name information to obtain the predicted system fault type;
[0009] Process the predicted system fault type, the fault scope information, and the expected recovery time according to a pre-generated fault library and an expert library to obtain fault information;
[0010] Perform fault injection on the system based on the fault information to conduct a self-healing evaluation of the system's faults.
[0011] Optionally, the process of using the fault type prediction model to process the system name information to obtain the predicted system fault type includes:
[0012] Query the system description information corresponding to the system name information from a CMDB (Configuration Management Database);
[0013] Preprocess the system description information to obtain preprocessed description information that meets the input requirements of the fault type prediction model;
[0014] Use the fault type prediction model to process the preprocessed description information to obtain the predicted system fault type.
[0015] Optionally, the process of processing the predicted system fault type, the fault scope information, and the expected recovery time according to a pre-generated fault library and an expert library to obtain fault information includes:
[0016] Based on the network topology structure of the resource data in the CMDB database and the fault scope information, screen out initial fault components from the pre-constructed fault library and expert library;
[0017] Based on the predicted system fault type, determine the target fault components among the initial fault components;
[0018] Generate the fault information according to the target fault components, the predicted system fault type, and the expected recovery time;
[0019] Among them, the fault information includes: target fault components, fault points, predicted system fault types, expected recovery times, and fault parameters.
[0020] Optionally, the process of determining the target fault components among the initial fault components based on the predicted system fault type includes:
[0021] Based on the predicted system fault type, determine the intermediate fault components among the initial fault components;
[0022] Obtain the target fault components selected by the user from the intermediate fault components.
[0023] Optionally, performing fault injection on the system based on the fault information to evaluate the fault self-healing of the system includes:
[0024] Generating a fault command according to the fault parsing result obtained by parsing the fault information;
[0025] When the fault injection node is a host node, injecting the fault command into the host node by using a preset binary tool;
[0026] When the fault injection node is a cluster node, injecting the fault command into the cluster node by using a pre-created custom resource.
[0027] In a second aspect, an embodiment of the present application provides a chaos engineering fault orchestration and injection device, and the device includes:
[0028] An information input module, configured to obtain system name information, fault range information, and an expected recovery time input by a user;
[0029] A prediction type obtaining module, configured to call a fault type prediction model to process the system name information to obtain a predicted system fault type;
[0030] A fault information obtaining module, configured to process the predicted system fault type, the fault range information, and the expected recovery time according to a pre-generated fault library and an expert library to obtain fault information;
[0031] A fault injection module, configured to perform fault injection on the system based on the fault information to evaluate the fault self-healing of the system.
[0032] Optionally, the prediction type obtaining module includes:
[0033] A description information obtaining unit, configured to query system description information corresponding to the system name information from a CMDB (Configuration Management Database);
[0034] An information processing unit, configured to preprocess the system description information to obtain preprocessed description information that meets the input requirements of the fault type prediction model;
[0035] A prediction type obtaining unit, configured to call the fault type prediction model to process the preprocessed description information to obtain the predicted system fault type.
[0036] Optionally, the fault information obtaining module includes:
[0037] An initial component screening unit, configured to screen out initial faulty components from a pre-constructed fault library and an expert library according to the network topology of the resource data in the CMDB database and the fault scope information;
[0038] A target component determination unit, configured to determine target faulty components among the initial faulty components based on the predicted system fault type;
[0039] A fault information generation unit, configured to generate the fault information according to the target faulty components, the predicted system fault type, and the expected recovery time;
[0040] Wherein, the fault information includes: target faulty components, fault points, predicted system fault types, expected recovery times, and fault parameters.
[0041] Optionally, the target component determination unit includes:
[0042] An intermediate component determination subunit, configured to determine intermediate faulty components among the initial faulty components based on the predicted system fault type;
[0043] A target component acquisition subunit, configured to acquire the target faulty components screened by the user from the intermediate faulty components.
[0044] Optionally, the fault injection module includes:
[0045] A fault command generation unit, configured to generate a fault command according to a fault parsing result obtained by parsing the fault information;
[0046] A first fault injection unit, configured to inject the fault command into the host node by using a preset binary tool when the fault injection node is a host node;
[0047] A second fault injection unit, configured to inject the fault command into the cluster node by using a pre-created custom resource when the fault injection node is a cluster node.
[0048] In a third aspect, an embodiment of the present application provides an electronic device, including:
[0049] A processor, a memory, and a computer program stored on the memory and executable on the processor, where the processor implements the chaos engineering fault orchestration and injection method described in any one of the above when executing the program.
[0050] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, when instructions in the storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the chaos engineering fault orchestration and injection method described in any one of the above.
[0051] Compared with the prior art, the embodiments of the present application include the following advantages:
[0052] In the embodiments of the present application, by obtaining the system name information, fault range information, and expected recovery time input by the user, calling the fault type prediction model to process the system name information, and obtaining the predicted system fault type. According to the pre-generated fault library and expert library, the predicted system fault type, fault range information, and expected recovery time are processed to obtain the fault information. Based on the fault information, fault injection is performed on the system to evaluate the fault self-healing of the system. By fully mining the knowledge information of the fault library and expert library, vectorizing and storing it, and training the large model, the embodiments of the present application can automatically arrange and inject faults by identifying the intention of fault injection input by the user, thereby reducing the intervention of manual scenario arrangement on the basis of covering various fault scenarios as much as possible.
[0053] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 is a flowchart of the steps of a chaos engineering fault orchestration and injection method provided by an embodiment of the present application;
[0055] Figure 2 is a schematic diagram of a fault injection process provided by an embodiment of the present application;
[0056] Figure 3 is a schematic diagram of the structure of a chaos engineering fault orchestration and injection device provided by an embodiment of the present application;
[0057] Figure 4 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0058] To make the above objects, features, and advantages of the present application more apparent and understandable, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0059] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. The singular forms "a", "said", and "the" used in the embodiments of the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.
[0060] Referring to Figure 1 , a flowchart of the steps of a chaos engineering fault orchestration and injection method provided by an embodiment of the present application is shown, as Figure 1As shown, the chaos engineering fault orchestration and injection method may include the following steps:
[0061] Step 101: Obtain the system name information, fault scope information, and expected recovery time input by the user.
[0062] Embodiments of this application can be applied to scenarios where, in combination with large models, the intention of fault injection input by the user is recognized, and fault orchestration and injection are automated.
[0063] Chaos engineering is a testing method that tests the resilience and recovery ability of a system by deliberately injecting faults into the system. It aims to discover defects that traditional testing methods may not be able to detect and improve the system's adaptability to real-world faults.
[0064] Large model: Refers to a machine learning model with a large number of parameters and a complex computational structure. These models are usually constructed by deep neural networks and have billions or even hundreds of billions of parameters. The design purpose of large models is to improve the model's expressive ability and prediction performance and be able to handle more complex tasks and data. Large models have a wide range of applications in various fields, including natural language processing, computer vision, speech recognition, and recommendation systems. Large models learn complex patterns and features by training on massive amounts of data and have stronger generalization ability, enabling accurate predictions for unseen data.
[0065] Fault injection is a reliability verification technique that deliberately introduces faults into the system through controlled experiments and observes the behavior of the system when faults exist. It is a fault verification technique that deliberately introduces faults into the system through controlled experiments and observes the behavior of the system when faults exist.
[0066] When performing chaos engineering fault orchestration and injection, the system name information, fault scope information, and expected recovery time input by the user can be obtained.
[0067] Among them, the system name information refers to the information input by the user for predicting the system fault type, which may include: information such as the system of an e-commerce trading platform, the system of a video playback type, etc.
[0068] The fault scope information refers to the fault scope, i.e., the blast radius, which refers to the scope of the host or container cluster, application service, and PaaS components affected by the fault. Usually, the blast radius is obtained through the network topology relationship drawn by CMDB resource data. For example, injecting a fault with an impact range of 3 that can be automatically recovered within 10 minutes in the XX e-commerce platform. Or, injecting a fault with an impact range of 2 that can be automatically recovered within 15 minutes in the XX video platform.
[0069] Understandably, the above examples are only examples listed for better understanding the technical solutions of the embodiments of the present application, and do not serve as the sole limitation to this embodiment.
[0070] A PaaS component refers to a software module or service unit that provides various basic services and functions for applications in a Platform as a Service (PaaS) architecture. They are the core components of the PaaS platform. By integrating and working together, they provide developers with a complete platform for developing, running, and managing applications. PaaS components usually include, but are not limited to, the aforementioned application servers, database management systems, message queues, cache systems, container orchestration tools, continuous integration / continuous deployment tools, log management systems, monitoring and alerting systems, etc. These components provide a series of functional supports for applications, such as computing resources, storage resources, network resources, data management, message passing, application deployment and management, log analysis, performance monitoring, etc., enabling developers to focus on implementing the business logic of the application without having to pay too much attention to the underlying infrastructure and operation and maintenance details.
[0071] CMDB (Configuration Management Database) is a database used to store and manage various configuration items and their relationships in an IT infrastructure, providing important data support for IT service management.
[0072] The expected recovery time refers to the expected time for a node with fault injection to perform fault recovery.
[0073] After obtaining the system name information, fault scope information, and expected recovery time input by the user, step 102 is executed.
[0074] Step 102: Invoke the fault type prediction model to process the system name information to obtain the predicted system fault type.
[0075] A fault type prediction model refers to a model used for predicting system fault types.
[0076] After obtaining the system name information, fault scope information, and expected recovery time input by the user, the fault type prediction model can be called to process the system name information to obtain the predicted system fault type. In a specific implementation, for the prompt information of the simulated fault scenario input by the user, the system description is obtained by querying the resource data in the CMDB database through the system name, and natural language processing is performed to identify the system type. Specifically, the system type can include: [CPU type / network type / memory type / disk type / combination type, etc.]+[data storage type / file storage type / cache type / retrieval type / combination type, etc.], etc. For these system types, a classifier based on the bert pre-trained model can be used for system type identification. For example, for a system with a function description of an e-commerce trading platform, for its high-concurrency spike scenario where faults are most likely to occur, it can be identified as a system type of [combination type]+[cache type]; for a system with a function description of a video playback type, it can be identified as a system type of [network type]+[file storage type].
[0077] The processing process of the fault type prediction model can be described in detail in combination with the following specific implementation manners.
[0078] In a specific implementation manner of the present application, step 102 above may include:
[0079] Sub-step S1: Query the system description information corresponding to the system name information from the CMDB (Configuration Management Database) database.
[0080] In this embodiment, after obtaining the system name information input by the user, the system description information corresponding to the system name information can be queried from the CMDB database. Specifically, a suitable database connection tool or library can be used (for example, using mysql-connector-python to connect to the MySQL database and using psycopg2 to connect to the PostgreSQL database in Python, etc.) to establish a connection with the CMDB database. Then, a query statement is constructed, that is, a query statement is constructed according to the system name information. For example, assume that there is a table named systems in the CMDB database, which contains fields such as system_name (system name) and system_description (system description). Finally, the query can be executed, that is, the above query statement is executed through the database connection, and the query result is obtained. The query result will contain the system description information corresponding to the input system name information.
[0081] After querying the system description information corresponding to the system name information from the CMDB database, perform sub-step S2.
[0082] Sub-step S2: Preprocess the system description information to obtain preprocessed description information that meets the input requirements of the fault type prediction model.
[0083] After obtaining the system description information, the system description information can be preprocessed to obtain preprocessed description information that meets the input requirements of the fault type prediction model. Specifically, preprocess the description input by the user. By adding [CLS] and [SEP] to mark the beginning and end of the sentence and [PAD] to pad the length, the text is converted into a format that meets the input of bert (i.e., the fault type prediction model in this example).
[0084] Of course, the preprocessing process of the system description information can also include: text cleaning (i.e., removing special characters, punctuation marks, HTML tags (if any) in the system description information), format conversion (such as converting all text to lowercase to ensure text consistency), stop word removal (removing common meaningless words (such as "the", "and", "is", etc.)), etc.
[0085] After preprocessing the system description information to obtain preprocessed description information that meets the input requirements of the fault type prediction model, perform sub-step S3.
[0086] Sub-step S3: Call the fault type prediction model to process the preprocessed description information to obtain the predicted system fault type.
[0087] After preprocessing the system description information to obtain preprocessed description information that meets the input requirements of the fault type prediction model, the fault type prediction model can be called to process the preprocessed description information to obtain the predicted system fault type. Specifically, the preprocessed description information can be input into the bert pre-trained model, and the pooled output at the [CLS] position is used as the sentence embedding and input to the linear layer Linear. The Softmax activation function is used for the output of Linear, and finally the predicted system fault type for this problem is output.
[0088] By querying the system description information from the CMDB database in the embodiments of the present application, detailed and accurate configuration and related information corresponding to the system name can be obtained, providing a rich data basis for subsequent analysis. This helps to clarify the specific characteristics of the system, component relationships, etc., making the prediction of fault types more targeted, avoiding blind guessing, and improving the accuracy of fault location. At the same time, preprocessing the system description information to meet the input requirements of the fault type prediction model can effectively reduce data noise and standardize the data format. The preprocessed data can be better understood and processed by the model, thereby improving the accuracy and stability of model prediction. For example, operations such as text cleaning, word segmentation, and vectorization can transform the original natural language text into numerical vectors that the model can process, enabling the model to more accurately capture the potential relationship between the key information in the text and the fault type. Moreover, by invoking the fault type prediction model to process the preprocessed description information, the learning ability and pattern recognition ability of the model can be utilized to automatically learn the mapping relationship between the system description and the fault type from a large amount of historical data and experience. Thus, intelligent analysis and prediction of new system description information can be realized, quickly and accurately giving the predicted system fault type, greatly improving the efficiency of fault diagnosis, reducing the time and cost of manual diagnosis, especially when facing complex systems and a large amount of fault data, the advantages of this automated prediction method are more obvious.
[0089] After obtaining the predicted system fault type, the user can interact with the large model system to annotate whether the type recognition is accurate, and iteratively update and train the model.
[0090] After invoking the fault type prediction model to process the system name information to obtain the predicted system fault type, step 103 is executed.
[0091] Step 103: Process the predicted system fault type, the fault range information, and the expected recovery time according to the pre-generated fault library and expert library to obtain fault information.
[0092] The fault library contains the summary information of the actual system faults that have occurred.
[0093] The expert library contains the fault information descriptions compiled by relevant professionals based on experience summaries.
[0094] Among them, the fault information description can use information such as fault components, fault points, fault types, influence ranges, and recovery times to describe the fault information:
[0095] Faulty components: basic resources (CPU, memory, network, disk, processes, files, system calls, etc.), docker containers (container CPU, container memory, container network, container disk, container processes, etc.), cloud-native platforms (pod basic resources, node basic resources, container basic resources, etc.), application services (host services, k8s containerized services), PaaS middleware (Redis, Mysql, Kafka, ES, etc.), etc.
[0096] Fault point: refers to the specific faulty component, such as a certain host (IP is XX.XX.XX.XX), a certain Mysql database (IP is XX.XX.XX.XX, port is XX), obtained through CMDB resource data, and generally INSTANCE-NAME is used to describe the fault point.
[0097] Fault type: specific fault classification injected for faulty components. Common fault types that can be injected by the chaosblade tool (an open-source chaos engineering experiment tool designed to help users conduct chaos engineering practices in production or test environments to improve system stability and reliability) are:
[0098] 1. Disk: IO (Input / Output) load, disk filling.
[0099] 2. Memory: memory occupancy.
[0100] 3. CPU: CPU (Central Processing Unit) load.
[0101] 4. Network: including network latency, network packet loss, network packet corruption, network packet duplication, network packet out-of-order, port occupancy, DNS (Domain Name System) resolution error.
[0102] 5. Application.
[0103] Fault parameters: for different fault types, there are different fault parameter inputs. For example: 1. For disk load faults, there are parameters such as disk directory, file block size, etc. 2. For application faults inside containers, there are parameters such as container ID, container name, namespace, pod name, etc.
[0104] Fault recovery time: refers to the duration from the start of the fault to the system recovery.
[0105] After obtaining the data from the expert database and the fault database, the data in the fault database and the expert database can be cleaned to obtain formatted fault information, which is usually represented in the form of json. The high-dimensional fault information description data is mapped to a low-dimensional space through Embedding and stored in the vector database Milvus. Among them, Milvus is a highly scalable open-source vector database that can manage a large amount of high-dimensional vector data and is very suitable for applications such as Retrieval Augmented Generation (RAG), semantic search, and recommendation systems.
[0106] Milvus can be integrated with mainstream Embedding models such as OpenAI and sentence-transformer through PyMilvus to generate Embedding vectors, as shown below:
[0107]
[0108] After constructing the fault database and the expert database, based on the pre-generated fault database and expert database, the predicted system fault type, fault range information, and expected recovery time can be processed to obtain fault information. The implementation process can be described in detail in combination with the following specific implementation methods.
[0109] In a specific implementation of the present application, the above step 103 may include:
[0110] Sub-step M1: According to the network topology structure of the resource data in the CMDB database and the fault range information, screen out the initial fault components from the pre-constructed fault database and expert database.
[0111] In this embodiment, after obtaining the fault range information input by the user, each component involved in the fault range information can be traversed. For each component, its associated information is searched in the network topology structure of the CMDB database to understand its location in the network and the other components it is connected to. Then, check whether the component exists in the fault database or the expert database. If it exists, it means that the component has the possibility of failing, and it is listed as an initial fault component. Until all components in the fault range information have been checked, an initial fault component list is finally obtained.
[0112] After screening out the initial fault components from the pre-constructed fault database and expert database according to the network topology structure of the resource data in the CMDB database and the fault range information, sub-step M2 is executed.
[0113] Sub-step M2: Based on the predicted system fault type, determine the target fault component among the initial fault components.
[0114] After screening out the initial faulty components, the target faulty components in the initial faulty components can be determined based on the predicted system fault type. Specifically, each component in the initial faulty component list can be traversed. For each component, the corresponding fault type of the component is searched in the fault library and the expert library. Furthermore, it is checked whether the predicted system fault type is included in these fault types. If it is included, it indicates that the component may be the source of the current fault, and it is determined as the target faulty component. This process continues until all components in the initial faulty component list have been checked, and finally a list of target faulty components is obtained.
[0115] After determining the target faulty components in the initial faulty components based on the predicted system fault type, sub-step M3 is executed.
[0116] Sub-step M3: Generate the fault information according to the target faulty component, the predicted system fault type, and the expected recovery time.
[0117] The fault information includes the target faulty component, the fault point, the predicted system fault type, the expected recovery time, and the fault parameters. Among them, the target faulty component is the component in the previously determined list of target faulty components, the predicted system fault type is the previously obtained prediction result, and the expected recovery time is the information input by the user.
[0118] After obtaining the target faulty components, the fault information can be generated according to the target faulty components, the predicted system fault type, and the expected recovery time. Specifically, for each target faulty component, by combining the network topology structure of the CMDB database and the information in the expert library, the specific location of the component in the network and the possible faulty locations are analyzed to determine the fault point. For example, if the target faulty component is a network device, the fault point may be a certain interface or a certain functional module of the device. The fault parameters need to be determined according to the predicted system fault type and the characteristics of the target faulty component. The historical data and experience in the fault library and the expert library can be referred to analyze the relevant parameters in similar fault situations. For example, if the predicted system fault type is network latency, the fault parameters may include the time range of the latency, the affected bandwidth, etc.
[0119] For each target faulty component, the target faulty component, the fault point, the predicted system fault type, the expected recovery time, and the fault parameters are combined into a fault information record. Furthermore, the fault information records of all target faulty components can be summarized to form the final fault information set.
[0120] In a specific implementation, the intermediate faulty components in the initial faulty components can be determined based on the predicted system fault type. And the target faulty components screened by the user from the intermediate faulty components are obtained. The following content is used as an example to illustrate this process:
[0121] Train the model through the fault library and the expert library, and combine the descriptions such as the importance level of each application in the CMDB library to locate the application most likely to have a fault; based on the impact scope input by the user, combine the network topology structure of the CMDB resource data to frame each PaaS component within the impact scope, and based on the identified system type, obtain the CMDB resources most likely to have a fault (including basic resources, PaaS components, application services, etc.) and give recommendations.
[0122] The user can select the components and CMDB resources for which fault injection needs to be performed from the recommended list, and the system automatically generates specific fault information:
[0123]
[0124] After processing the predicted system fault type, fault scope information, and expected recovery time according to the pre-generated fault library and expert library to obtain the fault information, perform step 104.
[0125] Step 104: Perform fault injection on the system based on the fault information to evaluate the fault self-healing of the system.
[0126] After processing the predicted system fault type, fault scope information, and expected recovery time according to the pre-generated fault library and expert library to obtain the fault information, fault injection can be performed on the system based on the fault information to evaluate the fault self-healing of the system. Specifically, after obtaining the fault information, the fault information can be parsed, and a fault command can be generated according to the fault parsing result. When the fault injection node (i.e., the node where fault injection needs to be performed, corresponding to the fault point in the above content) is a host node, the fault command is injected into the host node using a preset binary tool. When the fault injection node is a cluster node, the fault command is injected into the cluster node using a pre-created custom resource. For example, fault injection is performed through the open-source chaosblade tool of Alibaba, and different fault injection methods are used for hosts and k8s clusters: for hosts, the chaosblade binary tool is used for command injection, and for k8s clusters, the method of creating a custom resource CRD is used for fault injection. By parsing the generated fault information in json format, a fault command or a YAML file corresponding to the CRD resource can be automatically generated and distributed to the corresponding host or k8s cluster for fault injection, etc.
[0127] Fault self-healing verification process: Determine whether the system completes fault self-healing within the expected recovery time: If fault self-healing is successfully achieved, it is considered that the system operation meets the expectations; if the system does not recover within the expected time, this fault scenario needs to be focused on.
[0128] In the embodiments of the present application, by fully mining the knowledge information in the fault library and the expert library, vectorizing and storing it, and training the large model, the fault orchestration and injection are automatically performed by identifying the intention of the fault injection input by the user, so that the manual intervention in the scenario orchestration can be reduced on the basis of covering various fault scenarios as much as possible.
[0129] Next, the fault injection process will be described in detail in conjunction with Figure 2 As shown in Figure 2 The fault injection process may include:
[0130] 1. Fault information corpus sorting: The sources of fault information include the fault library and the expert library. The fault library contains the summary of the actual system faults that have occurred, and the expert library is the description of the fault information compiled by relevant professionals according to their experience, including information such as available fault components, fault points, fault types, impact scope, and recovery time.
[0131] 2. Vectorization of fault information: By cleaning the data in the fault library and the expert library, formatted fault information can be obtained, usually represented in the form of json. Through Embedding, the high-dimensional fault information description data is mapped to a low-dimensional space and stored in the vector library Milvus. Milvus can be integrated with mainstream Embedding models such as OpenAI and sentence-transformer through PyMilvus to generate Embedding vectors.
[0132] 3. Generate fault information through the operation and maintenance large model according to the basic information such as the system name, fault scope, and expected recovery time input by the user. Specifically, for the Prompt of the simulated fault scenario input by the user, the system description can be obtained by querying the resource data in the CMDB database through the system name, and the system type can be identified through natural language processing. The large model trains the model through the fault library and the expert library, and combines the descriptions of the importance levels of each application in the CMDB library to locate the application most likely to have a fault.
[0133] 4. Fault injection: It is carried out through the open-source chaosblade tool of Alibaba. Different fault injection methods will be adopted for the host and the k8s cluster. For the host, the chaosblade binary tool is used for command injection, and for the k8s cluster, the method of creating a custom resource CRD (Custom Resource Definition) is used for fault injection. By parsing the generated json-formatted fault information, the fault command or the YAML file corresponding to the CRD resource can be automatically generated and distributed to the corresponding host or k8s cluster for fault injection.
[0134] Finally, a self-healing check for faults can be performed, that is, to determine whether the system has completed self-healing of faults within the expected recovery time: if the self-healing of faults is successfully achieved, it is considered that the system is operating as expected. If the system does not recover within the expected time, this fault scenario needs to be focused on.
[0135] The chaos engineering fault orchestration and injection method provided by the embodiments of the present application obtains the system name information, fault scope information, and expected recovery time input by the user. Calls the fault type prediction model to process the system name information to obtain the predicted system fault type. Processes the predicted system fault type, fault scope information, and expected recovery time according to the pre-generated fault library and expert library to obtain fault information. Performs fault injection on the system based on the fault information to conduct a self-healing evaluation of the system. The embodiments of the present application fully mine the knowledge information in the fault library and expert library, perform vectorized storage and large model training, and automatically orchestrate and inject faults by identifying the intention of fault injection input by the user, thereby reducing the intervention of manual scenario orchestration on the basis of covering various fault scenarios as much as possible.
[0136] Refer to Figure 3 , which shows a schematic structural diagram of a chaos engineering fault orchestration and injection device provided by the embodiments of the present application. As Figure 3 shown, the chaos engineering fault orchestration and injection device 300 may include the following modules:
[0137] An information input module 310, configured to obtain the system name information, fault scope information, and expected recovery time input by the user;
[0138] A prediction type acquisition module 320, configured to call a fault type prediction model to process the system name information to obtain a predicted system fault type;
[0139] A fault information acquisition module 330, configured to process the predicted system fault type, the fault scope information, and the expected recovery time according to a pre-generated fault library and expert library to obtain fault information;
[0140] A fault injection module 340, configured to perform fault injection on the system based on the fault information to conduct a self-healing evaluation of the system.
[0141] Optionally, the prediction type acquisition module includes:
[0142] A description information acquisition unit, configured to query the system description information corresponding to the system name information from a CMDB (Configuration Management Database);
[0143] An information processing unit for preprocessing the system description information to obtain preprocessed description information that meets the input requirements of the fault type prediction model;
[0144] A prediction type acquisition unit for calling the fault type prediction model to process the preprocessed description information to obtain the predicted system fault type.
[0145] Optionally, the fault information acquisition module includes:
[0146] An initial component screening unit for screening out initial fault components from a pre - constructed fault library and an expert library according to the network topology of the resource data in the CMDB database and the fault scope information;
[0147] A target component determination unit for determining target fault components among the initial fault components based on the predicted system fault type;
[0148] A fault information generation unit for generating the fault information according to the target fault components, the predicted system fault type, and the expected recovery time;
[0149] Wherein, the fault information includes: target fault components, fault points, predicted system fault type, expected recovery time, and fault parameters.
[0150] Optionally, the target component determination unit includes:
[0151] An intermediate component determination subunit for determining intermediate fault components among the initial fault components based on the predicted system fault type;
[0152] A target component acquisition subunit for acquiring the target fault components selected by the user from the intermediate fault components.
[0153] Optionally, the fault injection module includes:
[0154] A fault command generation unit for generating a fault command according to the fault analysis result obtained by parsing the fault information;
[0155] A first fault injection unit for injecting the fault command into the host node using a preset binary tool when the fault injection node is a host node;
[0156] A second fault injection unit for injecting the fault command into the cluster node using a pre - created custom resource when the fault injection node is a cluster node.
[0157] The chaos engineering fault orchestration and injection device provided by the embodiments of the present application obtains the system name information, fault scope information, and expected recovery time input by the user. It calls the fault type prediction model to process the system name information to obtain the predicted system fault type. According to the pre-generated fault library and expert library, it processes the predicted system fault type, fault scope information, and expected recovery time to obtain fault information. Based on the fault information, it injects faults into the system to evaluate the fault self-healing of the system. The embodiments of the present application fully mine the knowledge information of the fault library and the expert library, perform vectorized storage and large model training, and automatically orchestrate and inject faults by identifying the intention of fault injection input by the user, so as to reduce the intervention of manual scenario orchestration on the basis of covering various fault scenarios as much as possible.
[0158] The embodiments of the present application also provide an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the computer program is executed by the processor, it implements the above-mentioned chaos engineering fault orchestration and injection method.
[0159] Figure 4 FIG. shows a schematic structural diagram of an electronic device 400 according to an embodiment of the present invention. As Figure 4 shown, the electronic device 400 includes a central processing unit (CPU) 401, which can execute various appropriate actions and processes according to the computer program instructions stored in the read-only memory (ROM) 402 or the computer program instructions loaded from the storage unit 408 into the random access memory (RAM) 403. In the RAM 403, various programs and data required for the operation of the electronic device 400 can also be stored. The CPU 401, ROM 402, and RAM 403 are connected to each other through a bus 404. The input / output (I / O) interface 405 is also connected to the bus 404.
[0160] A plurality of components in the electronic device 400 are connected to the I / O interface 405, including: an input unit 406, such as a keyboard, a mouse, a microphone, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a disk, an optical disc, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the electronic device 400 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0161] Each of the processes and treatments described above can be executed by the processing unit 401. For example, the method of any of the above embodiments can be implemented as a computer software program, which is tangibly included in a computer-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the CPU 401, one or more actions in the method described above can be performed.
[0162] Additionally, an embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the above-mentioned chaos engineering fault orchestration and injection method.
[0163] Each embodiment in this specification is described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other.
[0164] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a device, or a computer program product. Therefore, the embodiments of the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0165] The embodiments of the present application are described with reference to the flowcharts and / or block diagrams of the methods, terminals (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminals to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminals generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0166] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminals to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one process or multiple processes and / or blocksFigure 1 The functions specified in one or more boxes.
[0167] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal, so that a series of operation steps are executed on the computer or other programmable terminal to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal provide steps for implementing the functions specified in one Figure 1 one process or more processes and / or boxes Figure 1 step of the functions specified in one or more boxes.
[0168] Although the preferred embodiments of the embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concept. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present application.
[0169] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or terminal. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or terminal including the said element.
[0170] The above has introduced in detail a chaos engineering fault orchestration and injection method, a chaos engineering fault orchestration and injection device, an electronic device and a computer-readable storage medium provided by the present application. Specific examples are used in this article to elaborate the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A method for chaos engineering fault orchestration and injection, characterized in that The method includes: Obtaining system name information, fault range information, and expected recovery time input by a user; Invoking a fault type prediction model to process the system name information to obtain a predicted system fault type; Processing the predicted system fault type, the fault range information, and the expected recovery time according to a pre-generated fault library and expert library to obtain fault information; Performing fault injection on the system based on the fault information to perform a fault self-healing evaluation of the system.
2. The method according to claim 1, wherein The invoking a fault type prediction model to process the system name information to obtain a predicted system fault type includes: Querying system description information corresponding to the system name information from a CMDB (Configuration Management Database); Preprocessing the system description information to obtain preprocessed description information that meets the input requirements of the fault type prediction model; Invoking the fault type prediction model to process the preprocessed description information to obtain the predicted system fault type.
3. The method according to claim 1, wherein The processing the predicted system fault type, the fault range information, and the expected recovery time according to a pre-generated fault library and expert library to obtain fault information includes: Screening out initial fault components from a pre-constructed fault library and expert library according to the network topology structure of the resource data in the CMDB database and the fault range information; Determining a target fault component among the initial fault components based on the predicted system fault type; Generating the fault information according to the target fault component, the predicted system fault type, and the expected recovery time; Wherein, the fault information includes: a target fault component, a fault point, a predicted system fault type, an expected recovery time, and fault parameters.
4. The method according to claim 3, wherein The determining a target fault component among the initial fault components based on the predicted system fault type includes: Determining intermediate fault components among the initial fault components based on the predicted system fault type; Obtaining the target fault component selected by the user from the intermediate fault components.
5. The method according to claim 1, wherein The performing fault injection on the system based on the fault information to perform a fault self-healing evaluation of the system includes: Generating a fault command according to a fault parsing result obtained by parsing the fault information; When the fault injection node is a host node, injecting the fault command into the host node by using a preset binary tool; When the fault injection node is a cluster node, injecting the fault command into the cluster node by using a pre-created custom resource.
6. A chaos engineering fault orchestration and injection device, characterized in that, The device includes: An information input module, configured to obtain system name information, fault range information, and expected recovery time input by a user; A prediction type obtaining module, configured to invoke a fault type prediction model to process the system name information to obtain a predicted system fault type; A fault information obtaining module, configured to process the predicted system fault type, the fault range information, and the expected recovery time according to a pre-generated fault library and expert library to obtain fault information; A fault injection module for injecting faults into the system based on the fault information to evaluate the fault self-healing of the system.
7. The device according to claim 6, characterized in that, The prediction type acquisition module includes: A description information acquisition unit for querying the system description information corresponding to the system name information from a CMDB (Configuration Management Database); An information processing unit for preprocessing the system description information to obtain preprocessed description information that meets the input requirements of the fault type prediction model; A prediction type acquisition unit for calling the fault type prediction model to process the preprocessed description information to obtain the predicted system fault type.
8. The device according to claim 6, characterized in that The fault information acquisition module includes: An initial component screening unit for screening out initial fault components from a pre-constructed fault library and expert library according to the network topology structure of the resource data in the CMDB database and the fault scope information; A target component determination unit for determining target fault components among the initial fault components based on the predicted system fault type; A fault information generation unit for generating the fault information according to the target fault components, the predicted system fault type, and the expected recovery time; Wherein, the fault information includes: target fault components, fault points, predicted system fault types, expected recovery times, and fault parameters.
9. An electronic device, characterized in that, It includes: A processor, a memory, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the chaos engineering fault orchestration and injection method according to any one of claims 1 to 5.
10. A computer-readable storage medium, characterized in that, When the instructions in the storage medium are executed by the processor of the electronic device, the electronic device can execute the chaos engineering fault orchestration and injection method according to any one of claims 1 to 5.