Automated Cloud Application Remediation via Dependency Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In cloud environments, changes to one cloud-based application can cause bugs or problems in dependent applications, making it difficult and time-consuming to diagnose and resolve issues due to complex dependencies and the need for rollback of changes.
Innovation Solution
A system for automated application remediation based on change tickets that receives dependency indicators from APIs, generates a GUI for visualizing cloud-based applications, identifies change indicators associated with incident tickets, and transmits commands to rollback or rollforward changes based on dependency analysis and timing differences.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual analysis and rollback processes are used to diagnose and resolve issues in cloud-based applications, then reliability can be maintained, but loss of time and productivity decrease significantly
Solution Approach 1:
The system performs preliminary actions by automatically analyzing application dependencies and identifying problematic changes before manual intervention is needed. The automated remediation system proactively detects issues by comparing change indicators with incident tickets, determining causation relationships, and preparing rollback commands in advance, eliminating the time-consuming manual diagnosis process while maintaining reliability through systematic analysis
Solution Approach 2:
The system enables self-service by automatically resolving incidents without human intervention. The automated remediation system independently analyzes dependency indicators, identifies the root cause changes, executes rollback or rollforward commands, and updates incident tickets autonomously, allowing the system to service itself and dramatically reducing the time required for diagnosis and resolution while maintaining application reliability
2Productivity
If automated remediation is implemented to reduce diagnosis time, then productivity improves, but device complexity increases
Solution Approach 1:
The automated remediation system achieves multi-functionality by consolidating multiple incident resolution tasks into a single unified platform. It performs dependency analysis, change identification, causation determination, command generation, and ticket updating all through one system, improving productivity without proportionally increasing complexity as each function leverages the same core infrastructure and data models
Solution Approach 2:
The system uses an intermediary approach by introducing an automated remediation layer between the complex underlying cloud infrastructure and the incident management process. This intermediary automatically translates complex dependency relationships and change indicators into actionable rollback commands, hiding the complexity from users while maintaining high productivity through systematic automated analysis and execution
3Measurement precision
If comprehensive dependency analysis is performed to accurately identify problematic changes, then measurement precision improves, but computing resources are consumed
Solution Approach 1:
The system applies segmentation by dividing the comprehensive dependency analysis into manageable components. It separately processes dependency indicators from different sources, analyzes change indicators individually, and systematically evaluates causation relationships between specific changes and incidents. This segmented approach maintains high measurement precision in identifying problematic changes while reducing computing resource consumption by avoiding redundant full-system analyses
Solution Approach 2:
The system uses partial action by focusing analysis only on relevant portions of the system affected by incidents. Instead of continuously analyzing all dependencies, it triggers targeted analysis only when incidents occur, examining specifically the changes and dependencies related to those incidents. This approach achieves sufficient measurement precision for accurate change identification while significantly reducing overall computing resource consumption by avoiding excessive unnecessary analysis
Data Source
AI summary
In some implementations, a system may receive dependency indicators associated with a plurality of cloud-based applications and receive change indicators associated with changes to one or more first applications of the plurality of cloud-based applications. The system may receive an indicator associated with an incident ticket based on a problem with a second application of the plurality of cloud-based applications. The device may determine at least one of the change indicators associated with the incident ticket based on dependencies between the one or more first applications and the second application and based on a difference between a time associated with the incident ticket and a time associated with the at least one of the change indicators. The system may, based on determining the at least one of the change indicators, transmit a command to rollback at least one of the changes or to rollforward at least one change.


