Cloud Instance Allocation Using Spare Capacity and On-Demand Fallback
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Cloud service providers face inefficiencies due to excess computing resources sitting idle, while clients with fluctuating needs cannot afford service guarantees, leading to underutilization and increased costs.
Innovation Solution
A cloud resources allocation system (CAS) that maximizes resource usage by matching clients with spare computing capacity through on-demand and spare allocations, offering flexibility and cost savings by allowing clients to opt-in to spare instances with limited guarantees.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If cloud service providers allocate spare computing capacity to clients without service guarantees, then resource utilization increases, but service reliability deteriorates
Solution Approach 1:
The patent segments cloud computing resources into two distinct categories: subscribed-to instances with service level guarantees and spare instances without guarantees. This segmentation allows the system to allocate resources differently based on client needs and provider capacity, improving overall utilization while maintaining reliability for critical workloads.
Solution Approach 2:
The patent implements dynamic allocation where spare instances can be allocated to clients based on real-time provider capacity and client demand. The cloud service provider can flexibly grant or revoke access to spare instances based on current utilization levels, allowing resource utilization to increase when capacity is available while maintaining service reliability through guaranteed subscribed instances.
2Reliability
If cloud service providers maintain guaranteed uptime for subscribed instances, then service reliability improves, but resource utilization deteriorates due to idle spare capacity
Solution Approach 1:
The patent makes spare computing instances serve multiple functions: they remain available to fulfill service level agreements when needed, and simultaneously can be allocated to additional clients when provider capacity exceeds subscribed demand. This multi-functionality allows the same physical infrastructure to support both guaranteed service reliability and improved resource utilization.
Solution Approach 2:
The system allows cloud service providers to automatically allocate spare capacity to additional clients without impacting subscribed instances. The spare instances self-manage by being available when provider capacity is excess and automatically becoming unavailable when subscribed demand increases, eliminating the need for manual resource management while maintaining both reliability and utilization.
3Productivity
If cloud service providers allocate spare instances to multiple clients, then resource utilization improves, but allocation fairness deteriorates as instances may be revoked with limited notice
Solution Approach 1:
The patent requires clients to bid on or request spare instances in advance, and the cloud service provider establishes allocation rules beforehand. Clients are notified of allocation decisions and potential revocations according to predetermined policies, providing clarity and stability despite the dynamic nature of spare instance allocation.
Solution Approach 2:
The system implements feedback mechanisms where the cloud service provider monitors spare instance utilization and client needs, then adjusts allocations accordingly. Clients receive notifications about allocation status and potential revocations, allowing them to plan workloads appropriately. This feedback loop maintains allocation stability while maximizing resource utilization through transparent, rule-based management.
Data Source
AI summary
Disclosed herein are various embodiments for a cloud resources allocation system. An embodiment operates by determining that an application requests a plurality of instances to execute across one or more processors of a cloud services platform. One or more spare instances are requested to fulfill at least a portion of the plurality of instances, and an allocation of at least a subset of the requested one or more spare instances is received. An execution of the application is directed to the allocated subset of the one or more spare instances, including a first spare instance. A notification is received from the cloud services platform that the first spare instance is to be reallocated to a different process. A subsequent execution of the application is redirected to an on demand instance of the cloud services platform.


