Virtual Machine Scaling via Predictive Analytics
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current telecommunications network systems face inefficiencies due to over-provisioning, leading to high capital and operational expenditures, as they are designed to handle worst-case traffic scenarios, and existing scaling methods are often reactive and rely on heuristics, failing to adapt proactively to varying loads.
Innovation Solution
A method and system for proactive scaling of virtual machine instances based on real-time analytics, which generate dependency data and load metrics to determine the required number of VMs needed to meet performance metrics, using a service scaling description language and modules like data collectors, analytics, and rescale modules to dynamically adjust resource allocation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If network systems are designed to cover worst-case traffic scenarios with over-provisioning, then service level agreement (SLA) levels are fulfilled, but capital and operational expenditure increase significantly
Solution Approach 1:
The patent implements dynamic scaling of virtual machine instances based on real-time traffic conditions and predictive analytics. The system continuously monitors network traffic patterns and automatically adjusts the number of VM instances to match actual demand, transitioning from static over-provisioning to dynamic adaptive provisioning. This resolves the contradiction by maintaining SLA fulfillment through real-time responsiveness while eliminating unnecessary resource allocation during low-traffic periods.
Solution Approach 2:
The system employs predictive analytics and cross-correlation of network data to anticipate traffic surges before they occur. By analyzing historical traffic patterns and current trends, the system proactively scales resources in advance of predicted demand peaks, ensuring SLA compliance without requiring permanent over-provisioning. This preliminary action allows the network to be prepared for worst-case scenarios only when needed, rather than maintaining constant excess capacity.
2Adaptability or versatility
If reactive scaling based on system metrics is used, then resource allocation responds to current demand, but the system cannot proactively adapt to varying loads and traffic patterns
Solution Approach 1:
The patent implements a comprehensive feedback mechanism that combines real-time monitoring of system metrics with predictive analytics. The system continuously collects network traffic data, analyzes patterns through cross-correlation, and uses this feedback to both react to current conditions and predict future demands. This dual feedback loop enables the system to maintain high adaptability to current demand while simultaneously preparing for upcoming traffic variations, eliminating the time loss associated with purely reactive approaches.
3Reliability
If traditional provisioning methods are used, then network capacity is sufficient for peak loads, but resource utilization efficiency decreases due to idle capacity during average or typical cases
Solution Approach 1:
The patent dynamically changes the parameter of resource allocation based on predicted and actual traffic conditions. Instead of maintaining fixed provisioning levels, the system adjusts the number of active virtual machine instances in real-time according to traffic patterns and predictive analytics. This parameter change enables the network to maintain sufficient capacity for peak loads while optimizing resource utilization efficiency by scaling down during periods of lower demand, eliminating the waste of idle capacity.
Data Source
Figure 1
Figure 2
Figure 2
AI summary
A computer-implemented method, in a virtualized system comprising multiple virtual machine (VM) instances executing over physical hardware, for scaling one or more of the VM instances, the method comprising generating dependency data representing a causal relationship between respective ones of a set of VM instances indicating a direct and/or indirect dependency between the set of VM instances for a service or application to be executed by the set of VM instances, determining a current state of the set of VM instances, scaling one or more of the VM instances in the set in response to a predicted variation in system load using one or more scaling rules relating at least one load metric for the system to a predefined number of VM instances for one or more VM instances of the set to meet a performance metric for the service or application.