Enterprise Disruption Manager for Controlled Resiliency Testing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Enterprise systems face challenges with unreliable resiliency due to unexpected failures and complex interdependencies among hardware and software components, leading to poor user experiences, reduced quality of service, and unnecessary downtime.
Innovation Solution
A disruption manager is used to plan, schedule, and implement controlled disruptions in enterprise systems to test resiliency, identify weaknesses, and build a library of disruption events to improve system adaptability and usability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If controlled disruptions are implemented to test resiliency, then system reliability is improved, but system complexity increases
Solution Approach 1:
The patent implements preliminary action by proactively scheduling and executing controlled disruptions before actual failures occur. The disruption manager plans disruption events in advance, selecting optimal targets and timing to test system resiliency without causing uncontrolled damage. This allows the system to be strengthened preemptively rather than reactively.
Solution Approach 2:
The patent converts the harmful effect of disruptions into a beneficial testing mechanism. By intentionally introducing controlled disruptions and monitoring system responses, the system learns from these artificial failures to improve its resiliency. The harm of temporary service degradation is transformed into the benefit of enhanced system robustness and improved failure response capabilities.
2Adaptability or versatility
If disruptions are applied to test resiliency, then adaptability is improved, but user experience deteriorates
Solution Approach 1:
The disruption manager performs preliminary analysis to select optimal disruption timing that minimizes impact on user experience. By scheduling disruptions during periods of low system utilization or planned maintenance windows, the system can test resiliency while reducing the negative impact on end users. This preliminary planning allows adaptability testing without severe service degradation.
Solution Approach 2:
The patent applies partial disruptions rather than complete system failures. By targeting specific non-critical components or functions for disruption while maintaining core services, the system can test its adaptability and recovery mechanisms without completely degrading user experience. This selective approach allows learning from disruptions while preserving essential functionality.
3Measurement precision
If comprehensive disruption monitoring is implemented, then measurement precision is improved, but loss of time increases
Solution Approach 1:
The patent creates a virtual model or copy of the system's response to disruptions through monitoring and logging. Instead of requiring extensive manual analysis of each disruption event, the system captures responses in a structured format that can be analyzed automatically. This copying approach preserves measurement precision while reducing the time required for human review and analysis of disruption data.
Data Source
AI summary
Various embodiments are generally directed to techniques for utilizing disruptions to enterprise systems, such as to test and/or improve the ability of the enterprise system to recover from system failures, for instance. In many embodiments, an enterprise system may include two or more networked components, such as hardware components and software components. Some embodiments are particularly directed to generating a disruption scheme for an enterprise system based on analysis of one or more aspects of the enterprise system. For example, embodiments may include one or more of planning, scheduling, creating, timing, implementing, administering, and/or strengthening against a disruption to an enterprise system in a controlled and monitored manner. In many embodiments, administration of a disruption scheme may be monitored, recorded, and/or analyzed. In many such embodiments, a library of disruption events may be generated based on monitoring, recording, and/or analyzing implementation of the disruption scheme.


