Adaptable ML Models for Cloud Application Reliability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for managing application reliability in cloud service platforms are inefficient and lack automation, requiring manual intervention by Site Reliability Engineers, and are unable to adapt to the continuous changes and complexities of cloud environments, leading to challenges in maintaining high availability and handling unexpected events.
Innovation Solution
The implementation of adaptable machine learning models that automate IT operations such as self-healing, cloning, and recovery of applications, by identifying cloud components, determining their relationships, generating infrastructure as code, and using AI for verification and anomaly detection, enabling continuous learning and adaptive management of cloud infrastructure.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual management by SREs is used, then flexibility and adaptability are maintained, but productivity and efficiency decrease as scale increases
Solution Approach 1:
The system enables self-service through automated reliability management where the machine learning model autonomously monitors, detects anomalies, and executes recovery operations without requiring continuous manual intervention from SREs, thereby improving productivity while maintaining adaptability
Solution Approach 2:
The patent replaces manual mechanical operations with automated machine learning-based systems that use AI algorithms to analyze cloud infrastructure data, detect reliability issues, and execute recovery actions, substituting human expertise with intelligent automation
2Loss of time
If manual reliability management is implemented, then adaptability to changes is maintained, but time consumption increases
Solution Approach 1:
The system performs preliminary actions by continuously monitoring cloud infrastructure components and pre-identifying potential reliability issues before they impact application functionality, enabling proactive recovery actions that reduce overall time loss
Solution Approach 2:
The machine learning model operates continuously to monitor, detect, and respond to reliability issues without interruption, ensuring that useful actions are performed continuously rather than periodically, thereby reducing time loss from manual interventions
3Productivity
If cloud infrastructure is continuously automated, then productivity improves, but device complexity increases
Solution Approach 1:
The patent introduces a machine learning model as an intermediary layer between cloud infrastructure components and SREs, which automatically processes complex infrastructure data, detects anomalies, and executes recovery actions, thereby improving productivity while managing complexity through intelligent mediation
4Adaptability or versatility
If adaptability to cloud changes is increased, then reliability management effectiveness improves, but measurement and detection difficulty increases
Solution Approach 1:
The system replaces manual detection methods with machine learning-based anomaly detection that automatically identifies patterns and deviations in cloud infrastructure data, making adaptability to changes more effective while reducing the difficulty of detecting and measuring reliability issues
Data Source
AI summary
Techniques for automated application reliability management using adaptable machine learning models are disclosed. In one example, a computer-implemented method may include identifying cloud components of an application deployed in a cloud service platform, determining relationships between the cloud components of the application and between the cloud components and other applications to generate a plurality of sub-assemblies, determining dependencies among the plurality of sub-assemblies to generate a super-assembly, generating infrastructure as code for application cloud components of the super-assembly and the plurality of sub-assemblies using metadata of the application cloud components, and performing a management operation to create a cloud infrastructure of the application using the generated infrastructure as code and verifying reliability of the created application using an adaptable machine learning model.


