Automatic fault injection and analysis platform based on domestic server

By using an automated fault injection and detection system based on domestically produced servers, the problems of insufficient test coverage, complex scenario design, and poor compatibility of chaos testing tools on domestically produced servers have been solved. The system has achieved stable operation and efficient fault injection, improving the accuracy of test results and the reliability of the system.

CN121940288APending Publication Date: 2026-04-28FUJIAN YIRONG INFORMATION TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FUJIAN YIRONG INFORMATION TECH
Filing Date
2025-10-23
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing chaos testing tools suffer from insufficient test coverage, complex scenario design, limited real-time monitoring capabilities, and poor compatibility with different server architectures on domestically produced servers. This results in scattered test results, high analysis difficulty, high false positive rate, and high configuration cost.

Method used

Design an automated fault injection and detection system based on domestically produced servers, including an application service layer and an infrastructure layer. Through the collaborative work of the user center system, the scheduling center system and the infrastructure layer, the system achieves automation, real-time monitoring and compatibility of fault injection tasks. The system adopts standardized templates and visualization tools to simplify configuration, and combines a task orchestration mechanism to achieve automated execution of the fault injection process.

Benefits of technology

It enables stable operation of the system in various heterogeneous environments, reduces the complexity of user operations, improves test coverage and monitoring capabilities, reduces human error, and improves the reliability and robustness of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121940288A_ABST
    Figure CN121940288A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of chaos engineering, in particular to an automatic fault injection and analysis platform based on a domestic server. The system comprises an application service layer and an infrastructure layer. The application service layer comprises a user center system and a dispatching center system; the user center system is used for a user to configure and manage a fault injection task; the dispatching center system is used for receiving a task request of the user center, coordinating the infrastructure layer to execute a task and feeding back a task state to the user center at the same time; and the infrastructure layer comprises a server cluster for deploying a fault injection agent and a hardware detection component, and is used for executing a specific fault injection task under the control of the dispatching center. The invention aims to provide an automatic fault injection and detection system based on a domestic server so as to solve the problems of insufficient test coverage, complex scene design, limited real-time monitoring capability and poor compatibility of different server architectures.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of chaos engineering technology, and in particular to an automated fault injection and analysis platform based on domestically produced servers. Background Technology

[0002] In modern distributed systems and cloud-native applications, system stability and reliability are paramount. The stability and reliability of domestically produced servers have become key factors in ensuring business continuity. With the widespread adoption of cloud computing, microservice architecture, and containerization technologies, system complexity has increased significantly, and the probability and scope of failures have also expanded. To address the limitations of traditional testing methods in revealing anomalies, and to improve system resilience and robustness, chaos testing has emerged. First proposed by Netflix, it aims to test system stability and recovery capabilities by intentionally introducing failures. Chaos testing simulates failure scenarios to verify a system's ability to cope with abnormal conditions such as resource failures, service interruptions, and network latency, thereby helping enterprises identify and address potential weaknesses in system design.

[0003] Existing chaos testing has effectively improved the robustness of the system, but it still faces the following shortcomings in practical applications: Existing chaos testing tools often focus on specific testing methods and cannot fully cover the various failure scenarios that the system may encounter, resulting in limited test results.

[0004] The design of test scenarios, due to their diversity and complexity, already presents users with high configuration difficulty and maintenance costs. As distributed systems and cloud-native applications become increasingly complex, the design and configuration of test scenarios further exacerbates the complexity, leading to a significant increase in configuration difficulty and maintenance costs for users.

[0005] The test results generated by chaos testing tools come from different fault types and testing methods. The results are often scattered and difficult to unify, which increases the difficulty of analysis and interpretation, and is prone to misjudgment, thus affecting the accuracy of system health assessment.

[0006] In addition, compatibility issues with tools on different domestic server architectures may prevent tests from being executed properly in certain environments, thus affecting the reliability and consistency of the tests.

[0007] To address issues such as insufficient test coverage, complex scenario design, limited real-time monitoring capabilities, and poor compatibility with different server architectures, we designed an automated fault injection and detection system based on domestically produced servers to solve these technical problems. Summary of the Invention

[0008] To address the aforementioned issues, the present invention aims to provide an automated fault injection and detection system based on domestically produced servers, thereby resolving problems such as insufficient test coverage, complex scenario design, limited real-time monitoring capabilities, and poor compatibility with different server architectures.

[0009] To achieve the above objectives, the present invention adopts the following technical solution: including an application service layer and an infrastructure layer; The application service layer includes a user center system and a scheduling center system; The user center system is used for users to configure and manage fault injection tasks; The scheduling center system is used to receive task requests from the user center, coordinate the execution of tasks at the infrastructure layer, and simultaneously report the task status back to the user center. The infrastructure layer includes a server cluster that deploys fault injection agents and hardware detection components to execute specific fault injection tasks under the control of the scheduling center and to feed back the data during the execution process to the scheduling center system.

[0010] Furthermore, the user center system is specifically as follows: The user center system includes a fault management module, a task management module, and a risk management module; The fault management module is used to set the fault injection method, fault scenario, fault type and fault level, and provides basic configuration for fault injection tasks; The task management module is responsible for managing task orchestration, fault configuration, and monitoring indicators to ensure task distribution and monitoring during task execution. The risk management module is used to set the blast radius, risk isolation mechanism, and risk assessment, so as to assess and control the risk when injecting faults, record the exercise history of the exercise objects according to the exercise rules, and identify objects with exercise risks through data analysis, thereby promoting business parties to conduct regular exercises.

[0011] Furthermore, the dispatch center system is specifically as follows: Responsible for the overall coordination of task requests in the user center system, including task scheduling and distribution, to allocate fault injection tasks to appropriate server nodes for execution, ensuring that tasks can run efficiently; The system monitors task execution status in real time, collects task execution data in real time, ensures task execution progress, and determines whether there are any abnormalities in the task execution process. While reporting the task status, it sends recovery instructions through the scheduling center to restore the work, thereby ensuring the stability of the server cluster.

[0012] Furthermore, the infrastructure layer is specifically as follows: The specific host receives the probe injection task assigned by the scheduling center system, deploys the probe according to the system architecture, and performs fault drills according to the preset fault scenarios and configurations to test the server's fault response capabilities; it generates and collects the process data of the fault drills and returns it to the scheduling center for further analysis and improvement.

[0013] The present invention has the following beneficial effects: 1. This invention establishes a standardized template to comprehensively cover fault injection in various scenarios, while remaining compatible with different domestic server architectures, effectively reducing the complexity of user operations. This method ensures stable system operation in various heterogeneous environments and provides detailed analysis and monitoring of the machines involved in the fault injection process. Through the task orchestration module, users can easily configure and schedule multiple fault scenarios, automating the fault injection and recovery process, reducing the complexity and error risk of manual operations. Real-time monitoring and data acquisition help provide feedback on various system performance indicators when a fault occurs, helping users to promptly discover and repair potential system vulnerabilities and weaknesses, thereby significantly reducing the risk of faults in practical applications and improving the overall reliability of the system.

[0014] 2- The invention is based on fault injection probes, and a standardized test template has been designed and equipped with a visual design tool, which simplifies the process of creating complex test scenarios and reduces the need for manual intervention.

[0015] 3- This invention takes into account the compatibility of different server architectures and operating systems, and ensures the heterogeneous execution of fault injection tasks by automatically adapting the probe installation package.

[0016] 4. This invention, combined with a task orchestration mechanism and preset task templates, automates the fault injection process. Monitoring and data collection scripts perform real-time monitoring and data acquisition of the fault injection process, improving system efficiency. Attached Figure Description

[0017] Figure 1 This is a schematic diagram of the present invention; Figure 2 This is a flowchart illustrating the fault injection drill process of the present invention. Figure 3 The flowchart is a state machine diagram of the fault simulation task arrangement template of this invention. Detailed Implementation

[0018] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments: See Figures 1-2 As shown, the solution includes an application service layer and an infrastructure layer; The application service layer includes a user center system and a scheduling center system; The user center system is used for users to configure and manage fault injection tasks; The scheduling center system is used to receive task requests from the user center, coordinate the execution of tasks at the infrastructure layer, and simultaneously report the task status back to the user center. The infrastructure layer includes a server cluster that deploys fault injection agents and hardware detection components to execute specific fault injection tasks under the control of the scheduling center and to feed back the data during the execution process to the scheduling center system.

[0019] Furthermore, the user center system is specifically as follows: The user center system includes a fault management module, a task management module, and a risk management module; The fault management module is used to set the fault injection method, fault scenario, fault type and fault level, and provides basic configuration for fault injection tasks; The task management module is responsible for managing task orchestration, fault configuration, and monitoring indicators to ensure task distribution and monitoring during task execution. The risk management module is used to set the blast radius, risk isolation mechanism, and risk assessment, so as to assess and control the risk when injecting faults, record the exercise history of the exercise objects according to the exercise rules, and identify objects with exercise risks through data analysis, thereby promoting business parties to conduct regular exercises.

[0020] Furthermore, the dispatch center system is specifically as follows: Responsible for the overall coordination of task requests in the user center system, including task scheduling and distribution, to allocate fault injection tasks to appropriate server nodes for execution, ensuring that tasks can run efficiently; The system monitors task execution status in real time, collects task execution data in real time, ensures task execution progress, and determines whether there are any abnormalities in the task execution process. While reporting the task status, it sends recovery instructions through the scheduling center to restore the work, thereby ensuring the stability of the server cluster.

[0021] Furthermore, the infrastructure layer is specifically as follows: The specific host receives the probe injection task assigned by the scheduling center system, deploys the probe according to the system architecture, and performs fault drills according to the preset fault scenarios and configurations to test the server's fault response capabilities; it generates and collects the process data of the fault drills and returns it to the scheduling center for further analysis and improvement.

[0022] The working principle is roughly as follows: Figure 2As shown, the first step for the drill participants is to create a drill space within the system as the working environment for fault testing. Within this drill space, users can import a list of target hosts to be tested based on various fault orchestration task models and configure each host. Configuration includes selecting the type of fault to be injected (e.g., CPU usage, memory usage, disk I / O load, network latency) and setting relevant parameters (e.g., fault occurrence time, duration, fault intensity). The system will store these configurations and provide detailed operational guidance for subsequent drills. A single drill may execute multiple fault tasks. To achieve automated execution, this paper treats each task as an atomic operation, describes the drill tasks using a finite automaton model, and formulates a drill template, as follows. Figure 3 As shown, The state machine defines , , Multiple states represent different fault injection task states. Tasks flow sequentially according to a defined order. A new task can only begin after the preceding task has completed. Task orchestration can be performed in parallel based on the target host's load. and The task orchestration template will start running first. The task can be executed simultaneously after completion. and After the two tasks have finished executing, run... The mission concluded, and all tasks were completed, marking the end of the fault drill.

[0023] The dispatch center automatically distributes fault tasks to designated target hosts based on a pre-defined fault task orchestration model. It determines the timing, type, scope, and intensity of fault injection according to the exercise plan, and assigns specific task parameters to each target host based on pre-defined configurations, including fault type, fault intensity, and duration. Once task configuration is complete, the dispatch center distributes the task information to each target host (server) via the network, ensuring that each host receives the information and is ready to execute the fault injection operation. After successful task distribution, the system confirms the task reception status based on feedback from the host.

[0024] After the task is issued, the system uploads fault injection media to the target host, which consists of fault injection probe installation packages adapted to different operating systems and hardware architectures. These packages may contain multiple versions and configurations to support various target host environments. The system also performs basic verification on the user-uploaded installation packages to ensure their integrity and compatibility. The uploaded probe packages are stored in the system's resource library for later use. The system connects to the target host via SSH and selects the latest version matching the target host's architecture from the preset probe installation package library. Subsequently, the system downloads the installation package to the target host via file stream transfer.

[0025] After downloading, the system automatically executes the deployment process, including running the preset installation script, decompressing the installation package, configuring probe parameters, and completing the probe program installation. After installation, the system checks the probe's running status. If installation fails, the system logs error information and attempts to redeploy. The entire process covers multiple steps, including remote connection, file transfer, and installation configuration, to ensure the probe program can be correctly installed and run stably on the target host. The probe package can trigger fault injection and data collection on the target host. The probe package injects abnormal behaviors into the system through specialized scripts, such as simulating resource consumption (CPU, memory, disk) through process injection, simulating network latency by modifying system parameters, and forcibly stopping processes. Simultaneously, the probe package performs system status checks to ensure the target system is in an acceptable state and sets thresholds to prevent system crashes caused by injecting overly drastic faults. Based on continuous monitoring and data collection, the probe collects system performance metrics in real time during fault injection and records changes in these metrics before and after the fault occurs.

[0026] The target host parses the fault simulation information and initiates the simulation. Subsequently, the system downloads the monitoring and data collection script from the source server to the target host via file stream, ensuring the integrity and version compatibility of the plugins. Next, based on predefined fault types and parameter configurations, the system executes an automated script to enable the chaos probe package, performing a novel fault simulation on the target host. This process includes fault injection, resource stress testing, and other operations to simulate potential problems in a real-world environment.

[0027] Throughout the fault simulation exercise, the system synchronously launched monitoring and data acquisition scripts to continuously collect and record key performance indicators of the target host collected by the probes, such as CPU utilization, memory usage, disk activity, and network traffic. By analyzing this data in real time, the system can generate detailed reports, including information on the fault type, system response, and performance changes. These reports are promptly presented to the user, enabling them to quickly assess the exercise's results and make necessary adjustments and optimizations.

[0028] After the fault simulation exercise concludes, the system needs to restore the exercise environment to its pre-exercise normal state. The restoration process includes terminating the fault injection process, restoring occupied system resources, and restarting disabled services or processes to ensure the system returns to its pre-exercise healthy state. The system will automatically stop the probe program via a script and delete the probe installation files and data. If the probe generates temporary files or logs during execution, the system will clean these up to prevent them from affecting subsequent system operation.

[0029] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0030] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0031] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0032] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0033] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

Claims

1. An automated fault injection and analysis platform based on domestically produced servers, characterized in that: This includes the application service layer and the infrastructure layer; The application service layer includes a user center system and a scheduling center system; The user center system is used for users to configure and manage fault injection tasks; The scheduling center system is used to receive task requests from the user center, coordinate the execution of tasks at the infrastructure layer, and simultaneously report the task status back to the user center. The infrastructure layer includes a server cluster that deploys fault injection agents and hardware detection components to execute specific fault injection tasks under the control of the scheduling center and to feed back the data during the execution process to the scheduling center system.

2. The automated fault injection and analysis platform based on domestically produced servers according to claim 1, wherein the user center system is specifically as follows: The user center system includes a fault management module, a task management module, and a risk management module; The fault management module is used to set the fault injection method, fault scenario, fault type and fault level, and provides basic configuration for fault injection tasks; The task management module is responsible for managing task orchestration, fault configuration, and monitoring indicators to ensure task distribution and monitoring during task execution. The risk management module is used to set the blast radius, risk isolation mechanism, and risk assessment, so as to assess and control the risk when injecting faults, record the exercise history of the exercise objects according to the exercise rules, and identify objects with exercise risks through data analysis, thereby promoting business parties to conduct regular exercises.

3. The automated fault injection and analysis platform based on domestically produced servers according to claim 1, wherein the scheduling center system is specifically as follows: Responsible for the overall coordination of task requests in the user center system, including task scheduling and distribution, to allocate fault injection tasks to appropriate server nodes for execution, ensuring that tasks can run efficiently; The system monitors task execution status in real time, collects task execution data in real time, ensures task execution progress, and determines whether there are any abnormalities in the task execution process. While reporting the task status, it sends recovery instructions through the scheduling center to restore the work, thereby ensuring the stability of the server cluster.

4. The automated fault injection and analysis platform based on domestically produced servers according to claim 1, wherein the infrastructure layer specifically comprises the following: The specific host receives the probe injection task assigned by the scheduling center system, deploys the probe according to the system architecture, and performs fault drills according to the preset fault scenarios and configurations to test the server's fault response capabilities; it generates and collects the process data of the fault drills and returns it to the scheduling center for further analysis and improvement.