VM Configuration Recommendation via Bayesian Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The existing techniques for selecting optimal Virtual Machine (VM) hardware configurations for deep learning models are inefficient, manual, and do not adapt to changing cloud service providers, failing to consider model metrics like training time, error metrics, and carbon emissions, which are crucial for real-time performance optimization.

Innovation Solution

A system that uses a combination of mathematical modeling, clustering, benchmarking, and Bayesian optimization to generate recommendations for VM configurations based on historic data, user requirements, and cost functions, effectively addressing the dynamic nature of cloud services and DL model complexities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual techniques are used for selecting optimal VM hardware configuration, then human expertise can guide the selection process, but the process becomes inefficient and time-consuming

Engineering Contradiction:
Improveaccuracy of VM configuration selectionVSAvoidefficiency of VM configuration selection
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system performs self-service by automatically selecting optimal VM configurations through benchmarking and performance prediction algorithms, eliminating the need for manual human expertise while maintaining high accuracy in configuration selection

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical selection processes with automated computational systems that use benchmarking data and machine learning models to predict runtime performance, substituting human judgment with algorithmic decision-making

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If static solutions are used for VM configuration selection, then implementation is simpler, but the solutions do not adapt to changing hardware configurations across cloud service providers

Engineering Contradiction:
Improveadaptability to changing hardware configurationsVSAvoidcomplexity of configuration selection system
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system implements dynamic adaptability by continuously updating performance predictions based on changing hardware configurations and cloud service provider updates, allowing the VM selection process to adapt to new environments without manual reconfiguration

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent performs preliminary benchmarking and performance characterization of hardware configurations in advance, creating a knowledge base that enables rapid adaptation to changing cloud environments without requiring complex real-time analysis

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If comprehensive benchmarking is performed on all VM configurations, then accurate runtime performance prediction is achieved, but the process requires considerable amount of time

Engineering Contradiction:
Improveaccuracy of runtime performance predictionVSAvoidtime required for benchmarking
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system applies partial benchmarking by selecting only representative VM configurations for detailed testing, using clustering techniques to identify key hardware variants that capture the essential performance characteristics without benchmarking every possible configuration

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent creates universal performance models that can predict runtime behavior across different VM configurations based on limited benchmarking data, allowing the system to generalize from a small set of measured configurations to a large set of untested configurations

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Productivity

If existing solutions focus on cost minimization and resource utilization, then operational costs are reduced, but model metrics like training time and error metrics are not taken into consideration

Engineering Contradiction:
Improvetraining time optimizationVSAvoidsimplicity of existing solutions
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The system merges multiple optimization objectives including cost, training time, and model performance metrics into a unified VM selection framework, simultaneously considering economic and technical factors that existing separate solutions addressed independently

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20240012694A1System and method for recommending an optimal virtual machine (VM) instance
Publication Date: 2024.01.11 TATA CONSULTANCY SERVICES LTD
  • US20240012694A1 patent drawing
  • US20240012694A1 patent drawing
  • US20240012694A1 patent drawing

AI summary

This disclosure relates generally to recommending an optimal VM instance. The increased use of Deep Learning (DL) models in several domains has resulted in an increased demand for hardware configurations to enable heavy computations and faster performance to support the DL techniques. However, the identification of the optimal hardware configuration for the DL requirement is challenging and requires a considerable amount of time and expertise, considering the highly configurable model configuration of DL techniques. The disclosed optimal selection of VM comprises several techniques including benchmarking, using benchmarked results for building an approximation function and use a Bayesian Optimizer (BO) technique to iterate through the search space and generate recommendations of VM configurations, that effectively address the challenges arising due to the dynamic nature of cloud services—pricing and hardware configuration, large number of VM available across regions and cloud service providers and estimating for different types of training code.