Scenario-Based Fault Injection for Distributed System Resilience
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for determining the resilience of distributed software systems to failure scenarios are time-consuming and inefficient, as they require manual fault modeling and chaotic fault injection, which may not adequately cover all potential failure scenarios, especially those dependent on fault intensity and system topology changes.
Innovation Solution
A system and method that determines the topology of a distributed system to identify injection points for failure scenarios, prioritizes these scenarios based on historical data and system analysis, and injects them to assess the system's resilience, stopping if unexpected responses occur, and updates parameters based on feedback.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual fault modeling is used to determine system resilience, then comprehensive coverage of failure scenarios can be achieved, but significant time and effort are required
Solution Approach 1:
The system performs preliminary actions by automatically generating fault models based on historical outage data and system topology before resilience assessment begins. This pre-generation of fault scenarios eliminates the time-consuming manual fault modeling process while ensuring comprehensive coverage of potential failure scenarios.
Solution Approach 2:
The system serves itself by automatically generating and updating fault models using its own collected data from historical outages and current system topology. This self-service capability eliminates the need for manual intervention in fault model creation, significantly reducing the time required for resilience assessment while maintaining comprehensive scenario coverage.
2Loss of time
If chaotic fault injection is used to test system resilience, then time consumption is reduced, but coverage of all potential failure scenarios is insufficient
Solution Approach 1:
The system performs preliminary analysis of system topology and historical outages to generate a comprehensive set of relevant fault scenarios before injection begins. This ensures that all potential failure scenarios are covered in advance, eliminating the incomplete coverage problem of chaotic injection while maintaining efficiency through automated scenario generation.
Solution Approach 2:
The system dynamically adjusts fault injection parameters based on system topology and historical data, focusing injection efforts on the most relevant failure scenarios. This parameter optimization ensures comprehensive coverage of critical failure modes while reducing unnecessary injections, achieving both complete coverage and time efficiency.
3Reliability
If fault models are generated manually, then comprehensive failure scenarios can be identified, but the process becomes tedious and outdated when system changes occur
Solution Approach 1:
The system automatically generates and updates fault models by monitoring system topology changes and incorporating new outage data. This self-updating mechanism ensures fault models remain current with system changes without requiring manual redesign, eliminating the tedious maintenance process while maintaining comprehensive scenario identification.
Solution Approach 2:
The system continuously monitors system topology and outage data, using this feedback to automatically update fault models. This feedback loop ensures that fault models adapt to system changes in real-time, maintaining comprehensive failure scenario identification without manual intervention and reducing design complexity.
4Ease of operation
If only the number of machines is varied in fault injection, then simple testing is achieved, but fault intensity dependencies are not captured
Solution Approach 1:
The system automatically varies multiple fault parameters including fault intensity, duration, and type based on historical outage data and system characteristics. This multi-parameter approach captures the full range of failure scenarios including intensity-dependent failures while maintaining ease of operation through automated parameter selection and management.
Data Source
AI summary
A system determines a topology of a distributed system and determines, based on the topology, one or more injection points in the distributed system to inject failure scenarios. Each failure scenario including one or more faults and parameters for each of the faults. The system prioritizes the failure scenarios and injects a failure scenario from the prioritized failure scenarios into the distributed system via the one or more injection points. The system determines whether the injected failure scenario causes a response of the distributed system to fall below a predetermined level. The system determines resiliency of the distributed system to one or more faults in the injected failure scenario based on whether the injected failure scenario causes the response of the distributed system to fall below the predetermined level.


