Adaptive Latency Injection for Cloud Recovery Capacity Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for ensuring high-availability and disaster recovery in distributed computing environments are disruptive and fail to optimize resource utilization, leading to significant capacity overhead and increased costs.
Innovation Solution
Implementing adaptive latency injection by dynamically adjusting latency values based on request context to manage resource utilization, allowing non-disruptive operation and reducing the need for redundant capacity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional redundancy models (2.5N capacity) are used to ensure high availability and disaster recovery, then system reliability is improved, but capacity overhead and costs increase significantly
Solution Approach 1:
The system dynamically adjusts latency values based on current resource utilization conditions. When resources are under-utilized, higher latency is tolerated to reduce capacity needs. When resources are heavily utilized, latency is reduced to maintain performance. This dynamic adaptation allows the system to operate reliably with less than 2.5N capacity by flexibly trading off latency against capacity availability.
Solution Approach 2:
The patent changes the latency parameter adaptively based on system state. By modifying the latency parameter in response to resource utilization metrics, the system can maintain reliability with reduced capacity. The latency parameter becomes a controllable variable that adjusts to balance between service quality and resource consumption.
2Quantity of substance
If client-side rate-limiting, server-side throttling, or load shedding are implemented to reduce capacity usage, then capacity overhead is reduced, but service disruption and request failures occur
Solution Approach 1:
The patent converts the typically harmful effect of latency into a beneficial control mechanism. Instead of actively dropping or throttling requests (which causes service disruption), the system uses controlled latency as a pressure valve to smoothly regulate traffic flow. This transforms what is normally a performance degradation into a mechanism that maintains service continuity while reducing capacity needs.
Solution Approach 2:
The system implements periodic monitoring of resource utilization and adjusts latency values accordingly. This periodic adaptation allows the system to respond to changing conditions without causing service disruption, maintaining ease of operation while reducing capacity requirements.
3Reliability
If DR resources are provisioned for disaster recovery scenarios, then system reliability is improved, but resources remain idle for most of their lifespan increasing costs
Solution Approach 1:
The patent makes DR resources multi-functional by allowing them to serve both disaster recovery purposes and normal production workloads. During normal operations, DR resources handle regular traffic, keeping them active and useful. When disasters occur, these same resources immediately switch to recovery mode. This universality eliminates the idle time problem while maintaining reliability.
Solution Approach 2:
The system dynamically allocates resources between production and recovery functions based on current needs. DR resources are not statically dedicated but flexibly assigned, allowing them to be utilized during normal operations and automatically activated for disaster recovery when needed, thus eliminating waste while ensuring capability.
Data Source
AI summary
Embodiments of the disclosure provide systems and methods for reducing the capacity used to provide High Availability (HA) and Disaster Recovery (DR) in a distributed computing environment. According to one embodiment, dynamic recovery of a cloud-based resource can comprise setting a current latency value to an initial latency value and handling received requests with the current latency value. Current resource utilization can be detected while requests are being processed and a determination can be made as to whether the detected current resource utilization exceeds a predetermined threshold amount of resource utilization. In response to determining the detected current resource utilization does not exceed the threshold, the current latency amount can be maintained at the initial latency value. In response to determining the detected current resource utilization exceeds the threshold, the current latency value can be adjusted and injected into handling of received client requests.


