Fault drilling method, device, equipment, medium and product

Through the fault drill platform, the fault injection rule expressions are generated and split. The problem of fault injection consumes resources is solved, and accurate fault injection and business system stability is achieved.

CN120276895APending Publication Date: 2025-07-08CHINA UNITED NETWORK COMM GRP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510337718.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the prior art, the problem of failure injection consumes resources affects the call of resources by normal services.

Method used

The fault drill platform obtains drill tasks for the target service, generates intermediate fault injection rule expressions, and splits them, generates target fault injection rule expressions corresponding to each terminal SDK, and directly reuses the in-process resources of the business service for fault injection.

Benefits of technology

Accurate fault injection is achieved, reducing resource consumption and improving the reliability and stability of the business system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120276895A_ABST
    Figure CN120276895A_ABST
Patent Text Reader

Abstract

The invention provides a fault drilling method and device, equipment, a medium and a product. The method comprises the following steps: acquiring a drilling task for a target service; generating an intermediate fault injection rule expression according to task parameters carried by the drilling task and the initial fault injection rule expression; according to the terminal SDK of the target service, splitting the intermediate fault injection rule expression to generate a target fault injection rule expression corresponding to each terminal SDK; and sending the target fault injection rule expression corresponding to each terminal SDK to a target service, so that the target service realizes fault injection according to the target fault injection rule expression corresponding to each terminal SDK. The method solves the problem that fault injection consumes resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of chaos engineering, and in particular, to a fault drill method, device, equipment, medium and product. Background Art

[0002] With the rise of cloud computing, microservices architecture and large-scale distributed systems, software systems have become increasingly complex, making it difficult to predict the behavior of the system in the face of unexpected situations such as faults, network problems, resource limitations, etc. Chaos engineering discovers potential weaknesses in the system by actively introducing faults and uncertainties to simulate various abnormal situations in the real environment, and repairs these problems in advance to improve the reliability of the system.

[0003] In the prior art, the fault injection types of chaos engineering focus on resource layers such as the Central Processing Unit (CPU), memory, and network, and use the fault injection of resources to observe the changes in services, so as to evaluate the effect of the fault drill.

[0004] However, there is a problem that fault injection consumes resources in the prior art. Summary of the Invention

[0005] Embodiments of this application provide a fault drill method, device, equipment, medium and product to solve the problem that fault injection consumes resources.

[0006] In a first aspect, an embodiment of this application provides a fault drill method, which is applied to a fault drill platform and includes:

[0007] Obtain a drill task for a target service;

[0008] Generate an intermediate fault injection rule expression according to the task parameters carried in the drill task and the initial fault injection rule expression;

[0009] Split the intermediate fault injection rule expression according to the terminal SDK of the target service to generate a target fault injection rule expression corresponding to each terminal SDK;

[0010] Send the target fault injection rule expression corresponding to each terminal SDK to the target service, so that the target service implements fault injection according to the target fault injection rule expression corresponding to each terminal SDK.

[0011] In a possible implementation manner, the task parameters include: fault name, fault injection method, target service name, fault duration, attention index, and fault recovery condition.

[0012] In a possible implementation, an intermediate fault injection rule expression is generated according to the task parameters carried by the drill task and the initial fault injection rule expression, including:

[0013] Extract task parameters from the drill task;

[0014] Determine the initial fault injection rule expression corresponding to the target service according to the target service;

[0015] Write the task parameters into the corresponding positions in the initial fault injection rule expression to generate an intermediate fault injection rule expression.

[0016] In a possible implementation, the method further includes:

[0017] Receive the metric value of the metric concerned sent by the target service;

[0018] If the metric value of the metric concerned of the target service reaches the fault recovery condition, send a fault recovery instruction to the target service.

[0019] In a possible implementation, the method further includes:

[0020] Receive the metric value of the metric concerned sent by the dependent service, where the dependent service is another service that depends on the target service;

[0021] If the metric value of the metric concerned of the dependent service reaches the fault recovery condition, send a fault recovery instruction to the target service.

[0022] In a second aspect, an embodiment of the present application provides a fault drill method applied to a target service, including:

[0023] Obtain the target fault injection rule expression corresponding to each terminal software development kit (SDK) in the target service sent by the fault drill platform, where the target fault injection rule expression is generated based on the task parameters carried by the drill task and the initial fault injection rule expression;

[0024] Obtain communication context information according to the business data of each terminal SDK;

[0025] Execute the target fault injection rule expression according to the communication context information.

[0026] In a possible implementation, the task parameters include: fault name, fault injection method, target service name, fault duration, metric concerned, and fault recovery condition.

[0027] In a possible implementation, the method further includes:

[0028] Send the metric value of the metric concerned to the fault drill platform in real time.

[0029] In a possible implementation, the method further includes:

[0030] Receiving a fault recovery instruction sent by a fault drill platform;

[0031] Stopping the execution of the target fault injection rule expression.

[0032] In a third aspect, an embodiment of the present application provides a fault drill device, including:

[0033] A first acquisition module, configured to acquire a drill task for a target service;

[0034] A first generation module, configured to generate an intermediate fault injection rule expression according to the task parameters carried by the drill task and the initial fault injection rule expression;

[0035] A second generation module, configured to split the intermediate fault injection rule expression according to the terminal software development kit (SDK) of the target service to generate a target fault injection rule expression corresponding to each terminal SDK;

[0036] A sending module, configured to send the target fault injection rule expression corresponding to each terminal SDK to the target service, so that the target service implements fault injection according to the target fault injection rule expression corresponding to each terminal SDK.

[0037] In a fourth aspect, an embodiment of the present application provides a fault drill device, including:

[0038] A second acquisition module, configured to acquire the target fault injection rule expression corresponding to each terminal software development kit (SDK) in the target service sent by the fault drill platform, where the target fault injection rule expression is generated based on the task parameters carried by the drill task and the initial fault injection rule expression;

[0039] An obtaining module, configured to obtain communication context information according to the service data of each terminal SDK;

[0040] A processing module, configured to execute the target fault injection rule expression according to the communication context information.

[0041] In a fifth aspect, an embodiment of the present application provides an electronic device, including: a memory, a processor;

[0042] The memory stores computer execution instructions;

[0043] The processor executes the computer execution instructions stored in the memory, so that the processor executes the above first aspect and / or various possible implementation manners of the first aspect.

[0044] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, are used to implement the first aspect and / or various possible implementation manners of the first aspect as described above.

[0045] In a seventh aspect, an embodiment of the present application provides a computer program product including a computer program that, when executed by a processor, implements the first aspect and / or various possible implementation manners of the first aspect as described above.

[0046] A fault drill method, apparatus, device, medium, and product provided by an embodiment of the present application obtain a drill task for a target service through a fault drill platform, so as to understand the configuration of the drill task by the user; generate an intermediate fault injection rule expression according to the task parameters carried by the drill task and the initial fault injection rule expression, which is used to indicate the execution of the drill task; split the intermediate fault injection rule expression according to the terminal software development kit (SDK) of the target service to generate a target fault injection rule expression corresponding to each terminal SDK, so that the drill task can be accurately corresponded to each terminal SDK; send the target fault injection rule expression corresponding to each terminal SDK to the target service, so that the target service realizes customized fault injection according to the target fault injection rule expression corresponding to each terminal SDK, and solves the problem of resource consumption in fault injection. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0048] Figure 1 It is a schematic diagram of the scenario of the fault drill method provided by an embodiment of the present application;

[0049] Figure 2 It is a schematic flow chart of the fault drill method provided by an embodiment of the present application;

[0050] Figure 3 It is a schematic diagram of the inter-service dependency relationship provided by an embodiment of the present application;

[0051] Figure 4 It is a schematic structure diagram of the fault drill apparatus provided by an embodiment of the present application Figure 1 ;

[0052] Figure 5 It is a schematic structure diagram of the fault drill apparatus provided by an embodiment of the present application Figure 2 ;

[0053] Figure 6 It is a schematic structure diagram of the electronic device provided by an embodiment of the present application.

[0054] Through the above-mentioned accompanying drawings, specific embodiments of the present application have been shown, and more detailed descriptions will be provided hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed Description of Specific Embodiments

[0055] Here, exemplary embodiments will be described in detail, and examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0056] First, the terms related to the present application will be explained:

[0057] SDK: It refers to Software Development Kit, that is, a software development tool kit, a collection of tools for developing specific software or applications, which can provide developers with a convenient development environment, enabling them to focus on the development of core functions without having to build all basic functions from scratch;

[0058] Chaos engineering: It refers to a practical method of testing the resilience and recovery ability of a system by deliberately injecting faults into a distributed system, that is, by simulating fault scenarios in a real environment to discover the weaknesses of the system in advance and improve them;

[0059] MySQL: It refers to My Structured Query Language, that is, a relational database management system, suitable for various application scenarios, from simple applications to complex enterprise-level systems.

[0060] With the development of Internet technology, software systems have become more complex as the number of functions increases. In the face of huge business requirements, software systems need to maintain their own stability and maintainability, and chaos engineering is an important technology to meet this requirement, by actively introducing faults to discover system defects.

[0061] In the prior art, the fault injection logic can be encapsulated through an image to simulate the real business system calling cluster resources to achieve fault injection, so as to conduct fault drills under the condition of fault isolation. However, this way of fault injection additionally occupies the computing node resources required for normal business such as CPU, disk, and network in the cluster, and may affect the resource call of normal business due to the introduction of faults, resulting in the problem of resource consumption by fault injection.

[0062] In this regard, a fault drill method provided by an embodiment of the present application, in order to accurately perform a fault drill on a target service, obtains a drill task for the target service through a fault drill platform, so as to obtain task parameters related to the drill in the drill task. For each target service, there can be a corresponding initial fault injection rule expression to provide a template for the fault injection rule, so as to form an intermediate fault injection rule expression corresponding to the target task.

[0063] At the same time, in order to face different terminal SDKs under the target service, after splitting the intermediate fault injection rule expression, a target fault injection rule expression corresponding to each terminal SDK is generated and then sent to the target service. And in order to implement the fault drill only from the business level corresponding to the target service, the business data of each terminal SDK in the target service is obtained, and the communication context information containing communication details in the business data is used to combine with the target fault injection rule expression to implement fault injection, directly reusing the in-process resources of the business service. Therefore, only the terminal SDK is repackaged twice, the accurate injection of the fault task is realized, and the problem of resource consumption caused by fault injection is also solved.

[0064] Figure 1 It is a schematic diagram of the scenario of the fault drill method provided by an embodiment of the present application, as Figure 1 shown. The execution subject of this method can be a fault drill system. Among them, the fault drill system includes a client, a fault drill platform and a target service. The client can refer to an application or interface used by a user to configure a fault drill task, so as to input, configure and manage the drill task. The fault drill platform can refer to a platform responsible for planning, executing and monitoring fault drill activities, and executes specific drill operations according to instructions received from the client. The target service can refer to a service entity that actually accepts and responds to fault simulation, and executes corresponding actions according to instructions received from the fault drill platform.

[0065] This embodiment does not make special restrictions on the implementation manner of the execution subject, as long as the execution subject can obtain a drill task for the target service; generate an intermediate fault injection rule expression according to the task parameters carried by the drill task and the initial fault injection rule expression; split the intermediate fault injection rule expression according to the terminal SDK of the target service to generate a target fault injection rule expression corresponding to each terminal SDK; send the target fault injection rule expression corresponding to each terminal SDK to the target service, so that the target service can implement fault injection according to the target fault injection rule expression corresponding to each terminal SDK.

[0066] The technical solution of the present application and how the technical solution of the present application solves the above technical problems will be described in detail below with specific embodiments. The following several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0067] Figure 2 It is a schematic flowchart of a fault drill method provided by an embodiment of the present application. As Figure 2 shown, the method includes:

[0068] S201. The fault drill platform obtains a drill task for a target service.

[0069] Among them, the drill task may refer to an operation instruction set generated by the fault drill platform for simulating faults in the target service. The drill task can be entered by the user from the registration page provided by the fault drill platform, or the deployment information for different times and different target services can be collected uniformly and then entered.

[0070] In this step, the user can configure the set drill task on the fault drill platform, provide the configuration information of the drill task, can arrange the drill task temporarily, or plan to configure a large-scale drill task. Usually, the entire plan involved in the drill task needs to be approved by relevant personnel before fault injection can be performed on the service.

[0071] For example, a drill task of "Remote Dictionary Service (Redis) unavailable drill" can be configured on the fault drill platform, and the configuration information is: the Redis cluster starts to be unavailable at 9 pm, and observe the change in the number of requests processed per second (Queries Per Second, QPS) of a certain interface of application service C and application service A. When the QPS of service C or service A drops by a certain percentage, the availability of Redis is restored.

[0072] S202. The fault drill platform generates an intermediate fault injection rule expression according to the task parameters carried by the drill task and the initial fault injection rule expression.

[0073] Among them, the task parameters may refer to the configuration information related to the drill task, used to instruct the target service to execute the drill task. In a possible implementation manner, the task parameters may include a fault name, a fault injection method, a target service name, a fault duration, a concerned metric, and a fault recovery condition.

[0074] The fault name may refer to the type of fault simulation performed on the target service, such as network latency, or service unavailability, so as to distinguish different faults.

[0075] A fault injection method can refer to a method of introducing faults into a target service. There can be multiple fault injection methods. For example, software-based fault injection, by modifying memory data or control flow, or hardware-based fault injection, by changing hardware configuration parameters.

[0076] The target service name can refer to the name of the service for which the drill task needs to be executed. There can be multiple target services, and faults can be injected into the corresponding target services simultaneously or at different times.

[0077] The fault duration can refer to the length of time that a fault remains active after being injected, used to evaluate the ability of the target service to handle faults within a specific time, and to determine the recovery time of the target service to prevent the business supported by the target service from crashing.

[0078] The metrics to be monitored can refer to the key performance indicators or health indicators that need to be monitored during the fault drill. It can be the response time, the error rate, or the throughput, so as to measure the performance and recovery ability of the target service when facing faults.

[0079] The fault recovery condition can refer to the criteria for the target service to recover from the fault drill, so as to ensure the self-repair ability of the target service and the specified emergency plan.

[0080] It can be seen that the task parameters carried by the drill task can extract all the parameters in the drill task for the issuance of the fault task. For example, for the drill task of "Redis unavailable drill", the configuration content in the drill configuration platform includes the following information:

[0081] 1. Fault name: Redis connection timeout;

[0082] 2. Fault injection method: Redis cluster - slave xx

[0083] 3. Target service name: Service C;

[0084] 4. Fault duration: 5 minutes;

[0085] 5. Metrics to be monitored: Decrease in QPS of Service C;

[0086] 6. Fault recovery condition: Stop fault injection when the percentage decrease in the metrics to be monitored is 10%;

[0087] 7. Whether the fault is automatically recovered: Yes;

[0088] 8. Fault start time: xx year xx month xx day 21:00;

[0089] From the above example, a connection timeout failure can be simulated for a slave in the Redis cluster at 21:00 on xx / xx / xx for 5 minutes, and the impact on the service can be evaluated by observing whether the query QPS of service C decreases. When the QPS of service C decreases by 10%, the fault injection is automatically stopped and the system is restored to normal operation, so as to test and improve the stability and recovery ability of service C in the face of database connection problems.

[0090] The fault injection rule expression can refer to a set of commands indicating fault injection. The initial fault injection rule expression can refer to the basic logic or rules for introducing faults, and can be guided by a predefined template or script to implement fault simulation in the fault drill platform. In the embodiments of the present application, the initial fault injection rule expression is a rule description language for a rule engine, that is, the initial fault injection rule expression is a statement that can be recognized by the rule engine and meets the requirements of the rule engine for recognition. The type of the rule engine is not limited in the present application.

[0091] For example, a class named DisasterRecoveryDrillRequest is predefined. The startTime is used to record the time of fault injection and can represent the timestamp in long integer; the ResourceID represents the target resource identifier of the fault injection; the type describes the specific injection method of the fault; and isActive is a boolean value used to identify whether the drill request is in an active state.

[0092] The intermediate fault injection rule expression can refer to the initial fault injection rule expression combined with specific task parameters and is used to instruct the target service to execute the drill task.

[0093] It can be seen that this step shows that after obtaining the drill task in the fault drill platform, the task parameters in the drill task are extracted, and after combining the task parameters with the initial fault injection rule expression, an intermediate fault injection rule expression is obtained, which is used to instruct the target service to execute the drill task.

[0094] Optionally, the task parameters can also include the expected impact scope of the drill to ensure that the drill will not cause unnecessary interference or risks to irrelevant parts of the system. The expected impact scope includes all modules that directly interact with the target service, as well as the services and functional points that depend on these modules, especially the key business processes of the target service, so as to more comprehensively understand the impact of the fault on the overall service quality and user experience.

[0095] In a possible implementation, the task parameters can be extracted from the drill task, and then according to the target service, the corresponding initial fault injection rule expression of the target service is determined, and then the task parameters are written into the corresponding positions in the initial fault injection rule expression to generate the intermediate fault injection rule expression.

[0096] It can be seen that the generated initial fault injection rule expression varies with different target services, and separate rule files are generated for different resources. For example, an initial fault injection rule expression is determined according to the target service, and then according to the drill task requirements of target service C, task parameters are obtained and filled in the corresponding positions, and then the placeholder or default value in the initial fault injection rule expression is replaced, so as to generate an intermediate fault injection rule expression customized for target service C.

[0097] S203. The fault drill platform splits the intermediate fault injection rule expression according to the terminal software development kit (SDK) of the target service, and generates a target fault injection rule expression corresponding to each terminal SDK.

[0098] Among them, since the SDK provides the implementation of specific functions for business applications, such as message push, data analysis, payment and other functions. Business applications can quickly obtain these functions by integrating the SDK without having to develop from scratch. Therefore, the same business can include multiple SDKs to meet different business needs. Then, in the face of different SDKs, there will be fault injection rule expressions corresponding to different services.

[0099] Therefore, the target fault injection rule expression can refer to the fault injection rule expression for different terminal SDKs, which is used to perform fault injection on different target services.

[0100] S204. The fault drill platform sends the target fault injection rule expression corresponding to each terminal SDK to the target service, so that the target service can implement fault injection according to the target fault injection rule expression corresponding to each terminal SDK.

[0101] In this step, the fault drill platform sends the split intermediate fault injection rule expression, that is, the target fault injection rule expression, to the target service. After the target service receives and parses it, it can execute the fault drill task.

[0102] Optionally, the fault drill platform maintains a long connection with the target service to facilitate the real-time distribution of the target fault injection rule expression.

[0103] In a possible implementation manner, the fault drill platform receives the metric value of the metrics concerned sent by the target service; if the metric value of the metrics concerned of the target service reaches the fault recovery condition, a fault recovery instruction is sent to the target service.

[0104] It can be seen that the fault drill platform continuously monitors the real-time metric values of the metrics of interest sent by the target service. Once these metric values reach the pre-set fault recovery conditions, the fault drill platform will automatically identify this situation and send a fault recovery instruction to the target service, thereby ensuring that when the impact caused by the simulated fault reaches the expected threshold, the system can promptly recover from the fault state to avoid unnecessary further impacts or losses.

[0105] For example, for the drill task of "Redis unavailable drill", after the drill is started at 9 o'clock, it will continuously calculate whether the fault recovery conditions are met. When the fault recovery conditions are met, it will push the rule to stop the drill through the rule engine, and the corresponding execution engine will no longer execute the corresponding rule. When service C calls Redis, it is a normal service.

[0106] In a possible implementation, the fault drill platform receives the metric values of the metrics of interest sent by the dependent service, and the dependent service is another service that depends on the target service; if the metric values of the metrics of interest of the dependent service reach the fault recovery conditions, a fault recovery instruction will be sent to the target service.

[0107] Among them, the dependent service can refer to a service that depends on the target service to run. When the dependent service is affected by the simulated fault of the target service and the metrics of interest of them reach the pre-set fault recovery conditions, the fault drill platform will also automatically detect this situation. Once the fault recovery conditions are reached, the platform will send a fault recovery instruction to the target service to prompt the target service to resume normal operation.

[0108] For example, for the drill task of "Redis unavailable drill", when application service C directly depends on the service provided by the Redis cluster, and application service A depends on application service C, that is, application service A indirectly depends on the Redis cluster. Since Redis provides the function of interface data caching, it can improve the processing ability of the interface, which is manifested as improving the QPS of the interface. If the Redis cluster is unavailable, it will cause the QPS of service C and service A to drop. When the metric value of interest of application service A reaches the fault recovery conditions, a fault recovery instruction also needs to be sent to the target service to prompt the target service to resume normal operation and prevent further impacts on related services due to other services being affected by the fault drill.

[0109] S205. The target service obtains the target fault injection rule expression corresponding to each terminal software development tool (SDK) in the target service sent by the fault drill platform.

[0110] In this step, only the target task can receive the issued target fault injection rule expression. For example, for the drill task of "Redis Unavailability Drill", only the Redis SDK on which Service C depends will receive the instruction. The target fault injection rule expression is executed inside the SDK, and the defined rule is triggered at the defined time, so that Service C returns the error information configured correspondingly when calling Redis to obtain data.

[0111] S206. The target service obtains communication context information based on the business data of each terminal SDK.

[0112] Among them, the business data can refer to the information carried in the request from the client to the server and the data content returned by the server to the client, which is used to indicate the core information of specific business functions and service interactions. The business data includes parameters and results related to specific business logics, such as the content requested by the user, the user's operation behaviors (such as login, query, transaction, etc.), the user's authentication information, device details, application version number, etc.

[0113] The communication context information can refer to the communication details in the business data, which is used to comprehensively understand and process the interaction situation between services. The communication context information can be the protocol type. The protocols used for service - to - service communication supported by different SDKs determine how to parse the incoming and outgoing data. The communication context information can also be session information, such as session ID, creation time, and duration, which helps to track and manage the interaction process between a single user or device and the service. The communication context information can also be network attributes, including source IP address, target IP address, and port number, which are used to identify the specific locations and configurations of the two communication parties.

[0114] In this step, the business data provides the input and output information required to execute specific business logics, while the communication context information provides details and complete materials for understanding the environment and status of service interactions.

[0115] It can be seen that the solution of this application is not limited to a specific service architecture type or deployment environment. Whether it is a traditional monolithic service architecture or a modern microservice architecture, this solution can be applied. At the same time, the solution of this application is not restricted by physical machines, virtual machines or containerized deployment environments. The function of fault drill can be realized by upgrading the corresponding SDK or dependency package in the application system.

[0116] S207. The target service executes the target fault injection rule expression according to the communication context information.

[0117] In this step, when performing on-demand fault injection, it is necessary to match the communication context information with the target fault injection rule expression. Once the match is successful, the normal response process will be interfered with according to the configured fault mode to simulate specific types of faults, such as increasing latency, modifying return values, etc. This allows for the precise simulation of various fault situations for different service interaction scenarios, effectively testing the stability of the business system and the effectiveness of the coping strategies.

[0118] For example, for the drill task of "Redis Unavailability Drill", at 9 pm, each terminal software development kit (SDK) in the target service obtains the corresponding target fault injection rule expression, that is, service C is instructed to enter an unavailable state at the specified time. After receiving this instruction, when application service C attempts to call Redis, the SDK will make a judgment according to the preset rules, causing application service C to be unable to call.

[0119] In this case, after service C detects that Redis is unavailable, it will automatically degrade and switch to the backup logic to obtain the required data. Usually, the data in Redis is only used for caching. When service C fails to obtain data from Redis, it will obtain data through other means. This logical operation will increase the response time of service C and affect the QPS of service C. By monitoring the QPS of the interfaces between service C and service A, the impact of Redis unavailability on the performance of the entire business system can be effectively evaluated. This method not only helps the drill participants clearly observe the actual execution effect of the drill plan, but also enables them to deeply understand the specific impact of this simulated fault on the user experience and system response.

[0120] Optionally, service C passes user information when calling Redis, indicating that when drilling Redis unavailability, it can be configured from the user dimension. Users who have set specific rules will return an error message indicating Redis unavailability; while for users who have not set rules, the request will continue to be processed normally.

[0121] In a possible implementation, the target service sends the metric values of the metrics it is concerned about to the fault drill platform in real time, so that the fault drill platform can monitor the status of the target service in real time.

[0122] In a possible implementation, the target service receives the fault recovery instruction sent by the fault drill platform; and stops executing the target fault injection rule expression. It can be seen that when certain conditions are met, the target service needs to stop the fault drill to avoid affecting normal business, or stop when the drill plan is completed, so as to avoid harm to the business system.

[0123] A fault drill method provided by an embodiment of the present application obtains a drill task for a target service through a fault drill platform, generates an intermediate fault injection rule expression according to task parameters and an initial fault injection rule expression, and then splits and sends it to each terminal SDK in the target service, so that the target service can determine the status of each terminal SDK according to the service data of each terminal SDK to perform specific fault injection, realizing efficient and accurate fault injection for the target service. During this process, the platform continuously monitors the attention index values of the target service and dependent services, and issues a recovery instruction once the preset fault recovery condition is reached, ensuring the safety and controllability of the fault drill process and effectively improving the reliability and stability of the business system.

[0124] Figure 3 It is a schematic diagram of the inter-service dependency relationship provided by an embodiment of the present application, as Figure 3 shown, application service B and application service C are independent services, and application service A is a service that depends on application service C. The fault drill platform issues rules to application service B and application service C, affecting the communication process between application service B and MySQL, and affecting the communication process between application service C and Redis. Since application service A depends on application service C, the communication process between application service A and Redis will also be affected.

[0125] Figure 4 It is a schematic structure of the fault drill device provided by an embodiment of the present application Figure 1 as Figure 4 shown, the fault drill device 40 provided in this embodiment includes:

[0126] A first acquisition module 401, configured to acquire a drill task for a target service;

[0127] A first generation module 402, configured to generate an intermediate fault injection rule expression according to task parameters carried in the drill task and an initial fault injection rule expression;

[0128] A second generation module 403, configured to split the intermediate fault injection rule expression according to the terminal software development kit SDK of the target service to generate a target fault injection rule expression corresponding to each terminal SDK;

[0129] A sending module 404, configured to send the target fault injection rule expression corresponding to each terminal SDK to the target service, so that the target service realizes fault injection according to the target fault injection rule expression corresponding to each terminal SDK.

[0130] In a possible implementation manner, the first generation module 402 is further configured to extract task parameters from the drill task;

[0131] Determine the initial fault injection rule expression corresponding to the target service according to the target service;

[0132] Write the task parameters into the corresponding positions in the initial fault injection rule expression to generate an intermediate fault injection rule expression.

[0133] In a possible implementation, the sending module 404 is further configured to receive the metric value of the metrics concerned sent by the target service;

[0134] If the metric value of the metrics concerned of the target service reaches the fault recovery condition, send a fault recovery instruction to the target service.

[0135] In a possible implementation, the sending module 404 is further configured to receive the metric value of the metrics concerned sent by the dependent service, where the dependent service is another service that depends on the target service;

[0136] If the metric value of the metrics concerned of the dependent service reaches the fault recovery condition, send a fault recovery instruction to the target service.

[0137] The fault drill device provided in this embodiment can execute the method provided in the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.

[0138] Figure 5 The structural schematic of the fault drill device provided in the embodiment of the present application Figure 2 , as Figure 5 shown, the fault drill device 50 provided in this embodiment includes:

[0139] A second acquisition module 501, configured to acquire the target fault injection rule expression corresponding to each terminal software development kit (SDK) in the target service sent by the fault drill platform, where the target fault injection rule expression is generated based on the task parameters carried in the drill task and the initial fault injection rule expression;

[0140] An obtaining module 502, configured to obtain communication context information according to the service data of each terminal SDK;

[0141] A processing module 503, configured to execute the target fault injection rule expression according to the communication context information.

[0142] In a possible implementation, the processing module 503 is further configured to send the metric value of the metrics concerned to the fault drill platform in real time.

[0143] In a possible implementation, the processing module 503 is further configured to receive the fault recovery instruction sent by the fault drill platform;

[0144] Stop executing the target fault injection rule expression.

[0145] The fault drill device provided in this embodiment can execute the method provided in the above method embodiment, and its implementation principle and technical effect are similar. Details are not described herein in this embodiment.

[0146] Figure 6 It is a schematic structural diagram of an electronic device provided in an embodiment of the present application. As Figure 6 shown, the electronic device 60 provided in this embodiment includes: at least one processor 601 and a memory 602. Optionally, the device 60 further includes a communication component 603. Among them, the processor 601, the memory 602, and the communication component 603 are connected through a bus 604.

[0147] In a specific implementation process, at least one processor 601 executes computer-executable instructions stored in the memory 602, so that at least one processor 601 executes the above method.

[0148] For the specific implementation process of the processor 601, reference can be made to the above method embodiment, and its implementation principle and technical effect are similar. Details are not described herein again in this embodiment.

[0149] In the above embodiment, it should be understood that the processor may be a central processing unit (English: Central Processing Unit, abbreviated: CPU), or other general-purpose processors, digital signal processors (English: Digital Signal Processor, abbreviated: DSP), application-specific integrated circuits (English: Application Specific Integrated Circuit, abbreviated: ASIC), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules in the processor.

[0150] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (Non-volatile Memory, NVM), such as at least one disk memory.

[0151] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, the buses in the drawings of the present application are not limited to only one bus or one type of bus.

[0152] The present application also provides a computer program product, including a computer program, which implements the above method when executed by a processor.

[0153] The present application also provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the processor executes the computer-executable instructions, the above method is implemented.

[0154] The above-mentioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory (SRAM), an electrically erasable programmable read-only memory (EEPROM), an erasable programmable read-only memory (EPROM), a programmable read-only memory (PROM), a read-only memory (ROM), a magnetic memory, a flash memory, a magnetic disk or an optical disc. The readable storage medium can be any available medium accessible by a general-purpose or special-purpose computer.

[0155] An exemplary readable storage medium is coupled to the processor, so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an Application Specific Integrated Circuit (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in a device.

[0156] The division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be an indirect coupling or communication connection through some interfaces, devices or units, and can be in an electrical, mechanical or other form.

[0157] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0158] In addition, in each embodiment of the present invention, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0159] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0160] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the aforementioned storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0161] Finally, it should be noted that: After considering the specification and practicing the invention disclosed herein, those skilled in the art will easily think of other implementation schemes of the present invention. The present invention aims to cover any variations, uses, or adaptive changes of the present invention. These variations, uses, or adaptive changes follow the general principles of the present invention and include the common general knowledge or conventional technical means in the technical field not disclosed in the present invention. It is not limited to the exact structure described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

Claims

1. A fault drill method, characterized in that, Applied to a fault drill platform, the method includes: Obtain a drill task for a target service; Generate an intermediate fault injection rule expression according to the task parameters carried in the drill task and the initial fault injection rule expression; Split the intermediate fault injection rule expression according to the terminal software development kit (SDK) of the target service to generate a target fault injection rule expression corresponding to each terminal SDK; Send the target fault injection rule expression corresponding to each terminal SDK to the target service so that the target service implements fault injection according to the target fault injection rule expression corresponding to each terminal SDK.

2. The method according to claim 1, wherein The task parameters include: fault name, fault injection method, target service name, fault duration, attention indicators, and fault recovery conditions.

3. The method according to claim 1 or 2, wherein The generating an intermediate fault injection rule expression according to the task parameters carried in the drill task and the initial fault injection rule expression includes: Extract the task parameters from the drill task; Determine the initial fault injection rule expression corresponding to the target service according to the target service; Write the task parameters into the corresponding positions in the initial fault injection rule expression to generate the intermediate fault injection rule expression.

4. The method according to claim 2, wherein The method further includes: Receive the indicator value of the attention indicator sent by the target service; If the indicator value of the attention indicator of the target service reaches the fault recovery condition, send a fault recovery instruction to the target service.

5. The method according to claim 2, wherein The method further includes: Receive the indicator value of the attention indicator sent by a dependent service, where the dependent service is another service that depends on the target service; If the indicator value of the attention indicator of the dependent service reaches the fault recovery condition, send a fault recovery instruction to the target service.

6. A fault drill method, characterized in that Applied to a target service, the method includes: Obtain the target fault injection rule expression corresponding to each terminal software development tool (SDK) in the target service sent by the fault drill platform, where the target fault injection rule expression is generated based on the task parameters carried in the drill task and the initial fault injection rule expression; Obtain communication context information according to the business data of each terminal SDK; Execute the target fault injection rule expression according to the communication context information.

7. The method according to claim 6, characterized in that, The task parameters include: fault name, fault injection method, target service name, fault duration, attention indicators, and fault recovery conditions.

8. The method according to claim 7, wherein The method further includes: Send the indicator value of the attention indicator to the fault drill platform in real time.

9. The method according to any one of claims 6 - 8, characterized in that, The method further includes: Receive the fault recovery instruction sent by the fault drill platform; Stop executing the target fault injection rule expression.

10. A fault drill device, characterized in that, Includes: A first acquisition module for obtaining a drill task for a target service; A first generation module for generating an intermediate fault injection rule expression according to the task parameters carried in the drill task and the initial fault injection rule expression; A second generation module for splitting the intermediate fault injection rule expression according to the terminal software development kit (SDK) of the target service to generate a target fault injection rule expression corresponding to each terminal SDK; A sending module, configured to send the target fault injection rule expression corresponding to each terminal SDK to the target service, so that the target service implements fault injection according to the target fault injection rule expression corresponding to each terminal SDK.

11. A fault drill device, characterized in that, Comprising: A second obtaining module, configured to obtain the target fault injection rule expression corresponding to each terminal software development kit (SDK) in the target service sent by the fault drill platform, where the target fault injection rule expression is generated based on the task parameters carried in the drill task and the initial fault injection rule expression; A obtaining module, configured to obtain communication context information according to the service data of each terminal SDK; A processing module, configured to execute the target fault injection rule expression according to the communication context information.

12. An electronic device, characterized in that, Comprising: A memory and a processor; The memory stores computer execution instructions; The processor executes the computer execution instructions stored in the memory, so that the processor executes the method according to any one of claims 1-9.

13. A computer-readable storage medium, characterized in that, Computer execution instructions are stored in the computer-readable storage medium, and when the computer execution instructions are executed by a processor, they are used to implement the method according to any one of claims 1-9.

14. A computer program product, characterized in that, Comprising a computer program, which implements the method according to any one of claims 1-9 when executed by a processor.