GPU Cluster Allocation Using Model-Specific Power Profiles
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing cloud computing systems fail to optimize GPU resource allocation for machine learning tasks based on energy efficiency, leading to inefficient energy consumption due to varying energy requirements of different ML models and GPUs.
Innovation Solution
A method and apparatus for configuring a GPU cluster by measuring power consumption characteristics of each server for different models and using integer programming to assign servers such that the total power consumption is minimized, considering maximum throughput and model-specific energy consumption patterns.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If newer GPUs are used for all tasks, then energy efficiency is improved, but device complexity and cost increase
Solution Approach 1:
The patent applies local quality by assigning different GPU types to different ML models based on their specific energy consumption characteristics. Instead of using newer GPUs for all tasks, the system measures and analyzes power consumption for each GPU-model combination, then optimally assigns GPUs to models where they provide the best energy efficiency. This resolves the contradiction by making GPU allocation heterogeneous and model-specific rather than uniform.
2Productivity
If GPU allocation is based on performance prioritization, then productivity is improved, but energy efficiency deteriorates
Solution Approach 1:
The patent changes the allocation parameter from performance-based to energy consumption-based. Instead of allocating GPUs according to raw performance metrics, the system measures actual power consumption for each GPU-model combination and uses this energy parameter as the primary criterion for allocation. This resolves the contradiction by shifting the optimization objective from productivity to energy efficiency while maintaining adequate performance through the three conditions.
3Ease of operation
If GPU servers are assigned without considering model-specific energy profiles, then ease of operation is improved, but energy consumption increases
Solution Approach 1:
The patent applies preliminary action by measuring and storing power consumption characteristics for each GPU-model combination before actual task allocation. The system pre-establishes an energy consumption profile database that captures the specific energy requirements of different ML models on different GPUs. This resolves the contradiction by performing the complex measurement and analysis work in advance, making the actual allocation process simple while achieving optimal energy efficiency.
Data Source
AI summary
Provided is a method of configuring a cluster, which is a method of assigning graphics processing unit (GPU) servers in a cloud in which a plurality of machine learning (ML) services are executed using an apparatus for configuring a cluster. The apparatus for configuring a cluster is configured to measure the power consumption characteristics of each of the GPU servers constituting the cloud for each of a plurality of different models processing the plurality of ML services and assign at least one GPU server to each of the plurality of models using power consumption characteristics of each of the GPU servers for each of the plurality of models to configure a GPU cluster for each of the plurality of models.


