Machine Learning Capacity Reservations for Cloud Failover
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Manual capacity reservations for cloud-based computing instances lead to human errors, mismatches, and inefficiencies, risking unavailability during application failover and increasing costs.
Innovation Solution
A capacity reservation machine learning model that uses linear regression techniques to automate capacity reservations, analyzing current fleet trends and predicting necessary resources for failover, integrated with an automation workflow to align and visualize modifications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual capacity reservations are created, then human control and flexibility are maintained, but human errors and mismatches increase
Solution Approach 1:
The system performs self-service by automatically analyzing fleet trends and generating capacity reservations without human intervention. The machine learning model autonomously processes data, identifies patterns, and creates reservations, eliminating manual errors while maintaining operational flexibility through automated decision-making
Solution Approach 2:
The patent replaces the mechanical manual process of creating reservations with an automated machine learning system. The ML model substitutes human operators by analyzing fleet data, predicting capacity needs, and automatically generating reservations, thereby eliminating human errors while maintaining the ability to adapt to changing conditions
2Productivity
If manual reservations are altered one at a time, then precision in individual changes is maintained, but time consumption increases
Solution Approach 1:
The system performs preliminary action by proactively analyzing fleet trends and predicting future capacity needs before failures occur. The ML model continuously monitors fleet data and pre-generates appropriate capacity reservations, eliminating the need for reactive manual adjustments and significantly reducing the time required to ensure adequate capacity
Solution Approach 2:
The patent merges multiple individual reservation operations into a single automated process. Instead of altering reservations one at a time, the ML model analyzes the entire fleet portfolio and generates comprehensive capacity reservations in one operation, dramatically improving productivity while reducing the time loss associated with sequential manual changes
3Reliability
If ad-hoc capacity reservations are created, then flexibility in decision-making is maintained, but availability during failover decreases
Solution Approach 1:
The system implements feedback by continuously monitoring fleet performance data and using it to refine capacity reservation predictions. The ML model analyzes actual failover scenarios and utilization patterns, adjusting future reservations based on this feedback loop, thereby improving instance availability while managing complexity through data-driven automation
Solution Approach 2:
The patent applies parameter changes by dynamically adjusting capacity reservation parameters based on fleet trends and predicted needs. The ML model modifies reservation quantities, timing, and allocation parameters automatically, ensuring optimal instance availability during failover while managing complexity through systematic parameter optimization rather than ad-hoc decisions
4Reliability
If capacity reservations are increased to ensure failover coverage, then availability during failover is improved, but costs increase
Solution Approach 1:
The system applies partial action by generating capacity reservations that are precisely tailored to predicted needs rather than creating excessive reservations. The ML model analyzes fleet trends to determine the optimal partial capacity required for failover coverage, avoiding both insufficient reservations and wasteful over-provisioning, thereby improving failover coverage while maintaining cost efficiency
Solution Approach 2:
The patent implements dynamics by making capacity reservations adaptive and responsive to changing fleet conditions. The ML model continuously updates predictions based on current fleet trends, adjusting reservation levels dynamically to match actual needs, ensuring adequate failover coverage while optimizing resource utilization efficiency and avoiding static over-provisioning
Data Source
AI summary
Embodiments disclosed are directed to a computing system that performs operations for leveraging machine learning to automate capacity reservations for application failover in a cloud-based computing system. The computing system determines a simulated usage capacity of a set of applications executing in a first zone of a cloud-based computing system. The computing system then determines an amount of cloud-based computing instances in a second zone of the cloud-based computing system needed to maintain the simulated usage capacity in an event of a failover of the first zone. Subsequently, the computing system reserves the amount of cloud-based computing instances in the second zone.


