Cloud Controller SLA Migration for Network Latency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Centralized cloud computing systems face challenges in meeting strict delay and reliability requirements, and existing resource migration methods do not effectively preserve customer experience during migration, especially when network conditions change or load balances need to be adjusted.
Innovation Solution
A cloud controller system that monitors performance metrics, generates implicit Service Level Agreements (SLAs) based on customer experience, and migrates resources to maintain or improve performance by selecting devices that meet both explicit and implicit SLA requirements, ensuring minimal disruption to the customer's experience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If resources are migrated to improve load balance or network conditions, then resource utilization and network performance are improved, but customer experience and SLA compliance may deteriorate during migration
Solution Approach 1:
The system performs preliminary actions by establishing implicit SLAs based on historical performance data before migration occurs. This allows the system to predict and maintain customer experience levels during migration by pre-defining performance thresholds that must be preserved, thus resolving the contradiction between improving resource utilization and maintaining SLA compliance.
Solution Approach 2:
The system continuously monitors performance metrics and uses feedback loops to adjust migration decisions. By comparing actual performance against implicit SLAs during migration, the system can dynamically control the migration process to ensure SLA compliance is maintained while still achieving resource utilization improvements.
2Loss of time
If distributed data centers are used to reduce network propagation delay, then latency is reduced, but system complexity increases
Solution Approach 1:
The cloud controller acts as an intermediary that manages the complexity of distributed data centers. It handles resource allocation, performance monitoring, and migration coordination across multiple geographically distributed locations, thereby reducing network propagation delay for customers while centralizing the complexity management function.
Solution Approach 2:
The system segments cloud services across multiple distributed data centers, placing resources geographically closer to customers to reduce network propagation delay. Each data center operates semi-independently under centralized coordination, allowing the system to leverage distribution benefits while managing complexity through modular architecture.
3Adaptability or versatility
If resources are moved to meet changing network conditions or load requirements, then network performance and load balance are improved, but customer experience may be disrupted
Solution Approach 1:
The system establishes implicit SLAs based on historical performance data before migration, creating a baseline for customer experience. This preliminary action allows the system to adapt to changing network conditions and load requirements while maintaining customer experience by ensuring migrations only occur when implicit SLA thresholds are met.
Solution Approach 2:
The system performs preliminary anti-action by identifying and preventing migration scenarios that would violate implicit SLAs. Before executing migrations to adapt to network conditions or load changes, the system checks against established performance baselines and blocks migrations that would disrupt customer experience, thus counteracting potential harm in advance.
Data Source
Figure 1
Figure 2~3
Figure 4~6
AI summary
Various exemplary embodiments relate to a method and related network node including one or more of the following: receiving performance metrics from at least one device within the cloud network; analyzing the performance metrics to generate at least one application requirement, wherein the at least one application requirement indicates a recent performance associated with a resource provisioned for a customer, wherein the resource is supported by a first device within the cloud network; identifying a second device within the cloud network to support the resource based on the at least one application requirement; and migrating the resource from the first device to the second device. The resource can be a virtual machine (VM) or a storage.