Scenario-Based Fault Injection for Distributed System Resilience

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for determining the resilience of distributed software systems to failure scenarios are time-consuming and inefficient, as they require manual fault modeling and chaotic fault injection, which may not adequately cover all potential failure scenarios, especially those dependent on fault intensity and system topology changes.

Innovation Solution

A system and method that determines the topology of a distributed system to identify injection points for failure scenarios, prioritizes these scenarios based on historical data and system analysis, and injects them to assess the system's resilience, stopping if unexpected responses occur, and updates parameters based on feedback.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual fault modeling is used to determine system resilience, then comprehensive coverage of failure scenarios can be achieved, but significant time and effort are required

Engineering Contradiction:
Improvesystem resilience assessmentVSAvoidtime to verify failure scenarios
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by automatically generating fault models based on historical outage data and system topology before resilience assessment begins. This pre-generation of fault scenarios eliminates the time-consuming manual fault modeling process while ensuring comprehensive coverage of potential failure scenarios.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system serves itself by automatically generating and updating fault models using its own collected data from historical outages and current system topology. This self-service capability eliminates the need for manual intervention in fault model creation, significantly reducing the time required for resilience assessment while maintaining comprehensive scenario coverage.

Inventive Principle:
Principle #25Self-service

2Loss of time

If chaotic fault injection is used to test system resilience, then time consumption is reduced, but coverage of all potential failure scenarios is insufficient

Engineering Contradiction:
Improvetime to assess resilienceVSAvoidcompleteness of failure scenario coverage
Core Design Contradiction:
Loss of timeVSReliability

Solution Approach 1:

The system performs preliminary analysis of system topology and historical outages to generate a comprehensive set of relevant fault scenarios before injection begins. This ensures that all potential failure scenarios are covered in advance, eliminating the incomplete coverage problem of chaotic injection while maintaining efficiency through automated scenario generation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system dynamically adjusts fault injection parameters based on system topology and historical data, focusing injection efforts on the most relevant failure scenarios. This parameter optimization ensures comprehensive coverage of critical failure modes while reducing unnecessary injections, achieving both complete coverage and time efficiency.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If fault models are generated manually, then comprehensive failure scenarios can be identified, but the process becomes tedious and outdated when system changes occur

Engineering Contradiction:
Improvefailure scenario identificationVSAvoidfault model design complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system automatically generates and updates fault models by monitoring system topology changes and incorporating new outage data. This self-updating mechanism ensures fault models remain current with system changes without requiring manual redesign, eliminating the tedious maintenance process while maintaining comprehensive scenario identification.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system continuously monitors system topology and outage data, using this feedback to automatically update fault models. This feedback loop ensures that fault models adapt to system changes in real-time, maintaining comprehensive failure scenario identification without manual intervention and reducing design complexity.

Inventive Principle:
Principle #23Feedback

4Ease of operation

If only the number of machines is varied in fault injection, then simple testing is achieved, but fault intensity dependencies are not captured

Engineering Contradiction:
Improvefault injection simplicityVSAvoidfailure scenario detection accuracy
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The system automatically varies multiple fault parameters including fault intensity, duration, and type based on historical outage data and system characteristics. This multi-parameter approach captures the full range of failure scenarios including intensity-dependent failures while maintaining ease of operation through automated parameter selection and management.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10467126B2Scenarios based fault injection
Publication Date: 2019.11.05 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10467126B2 patent drawing
  • US10467126B2 patent drawing
  • US10467126B2 patent drawing

AI summary

A system determines a topology of a distributed system and determines, based on the topology, one or more injection points in the distributed system to inject failure scenarios. Each failure scenario including one or more faults and parameters for each of the faults. The system prioritizes the failure scenarios and injects a failure scenario from the prioritized failure scenarios into the distributed system via the one or more injection points. The system determines whether the injected failure scenario causes a response of the distributed system to fall below a predetermined level. The system determines resiliency of the distributed system to one or more faults in the injected failure scenario based on whether the injected failure scenario causes the response of the distributed system to fall below the predetermined level.