Dynamic Resilience Optimization for Network Process Interdependencies
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technology remediation systems face challenges in identifying the root cause of issues in complex computer systems, particularly in cloud-based environments with numerous interdependencies, leading to inefficiencies in mean time to detect (MTTD) and mean time to repair (MTTR), and mean time between failures (MTBF).
Innovation Solution
A method and system that monitor computer-implemented processes, generate pre-breakage and simulation result snapshots with robustness and risk scores, and utilize a rules engine to simulate break events and apply response strategies, optimizing network architecture to reduce risk and enhance robustness.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If commercial systems require specific break events to trigger a single response, then the response is simple and direct, but the system cannot handle complex interdependencies in cloud-based environments
Solution Approach 1:
The patent segments the complex remediation process into multiple independent analysis components, each evaluating specific aspects of the break event. Multiple responses are generated and evaluated separately, then combined through a scoring mechanism to determine the optimal remediation strategy, enabling the system to handle complex interdependencies while maintaining manageable complexity
Solution Approach 2:
The system dynamically adjusts the remediation approach by evaluating multiple potential responses and selecting the optimal one based on real-time system state and interdependency analysis. The response selection is not static but adapts to the specific complex scenario being faced, allowing the system to handle varying degrees of complexity
2Loss of time
If humans perform production system support and incident triage, then the system can handle complex situations, but the response time is slow
Solution Approach 1:
The system performs self-service by automatically analyzing break events, evaluating multiple responses, and selecting optimal remediation strategies without requiring manual human intervention. The automated analysis components and scoring mechanisms enable the system to handle complex situations independently, dramatically reducing response time while maintaining the capability to address intricate interdependencies
3Reliability
If chaos testing systems interject disruptions like CPU spikes or network cutoff, then the system can test resilience, but the test scenarios are unrealistic representations of actual system behavior
Solution Approach 1:
The system performs preliminary analysis of actual system interdependencies and break event patterns before testing resilience. By pre-configuring test scenarios based on real system behavior and relationships, the chaos testing becomes more realistic while still maintaining the ability to stress-test resilience, bridging the gap between theoretical disruptions and actual system responses
Data Source
AI summary
Disclosed are hardware and techniques for testing computer processes in a network system by simulating computer process faults and identifying risk associated with correcting the simulated fault and identifying computer processes that may depend on the corrected computer process. The interdependent computer processes in a network may be determined by evaluating a risk matrix having a risk score and non-functional requirement score. An analysis of the risk score and non-functional requirement score accounts for interdependencies between computer processes and identified corrective actions that may be used to determine an optimal network environment. The optimal network environment may be updated dynamically based on changing computer process interdependencies and the determined risk and robustness scores.


