Asymmetric CPU GPU Share Allocation for Cloud VMs

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional solutions for implementing virtual machines in cloud computing fail to optimize the processing capabilities of allocated CPUs and GPUs, leading to inefficient usage and higher costs, especially due to the significant price difference between CPU and GPU resources.

Innovation Solution

The method involves asymmetrically sharing the processing resources of CPUs and GPUs among virtual machines or software instances, where the number of GPU shares is greater than the number of CPU shares, allowing for dynamic allocation based on computing demands and reducing the average cost of provisioning GPU resources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional solutions allocate CPU and GPU shares equally to virtual machines, then the allocation is simple and fair, but the processing capabilities are not utilized optimally resulting in inefficient usage and higher costs

Engineering Contradiction:
Improveprocessing capability utilizationVSAvoidallocation mechanism complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent applies local quality by differentiating the allocation of CPU and GPU shares based on the specific processing requirements of different virtual machines. Instead of uniform allocation, the system assigns customized share ratios (e.g., 80:20, 70:30, 60:40) tailored to each VM's workload characteristics, allowing GPU-intensive applications to receive more GPU shares while CPU-intensive applications receive more CPU shares, thereby optimizing processing capability utilization.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent implements dynamics by enabling flexible and adaptable allocation of CPU and GPU shares based on varying computing demands. The system allows dynamic adjustment of share ratios in response to changing workload patterns, allowing the allocation mechanism to evolve from static equal sharing to dynamic demand-based distribution, improving productivity while managing complexity through automation.

Inventive Principle:
Principle #15Dynamics

2Reliability

If GPU resources are allocated at higher prices due to scarcity and higher performance requirements, then the processing performance is improved, but the overall cost of cloud computing services increases

Engineering Contradiction:
Improveprocessing performanceVSAvoidcost of processing resources
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent applies parameter changes by varying the share ratios of CPU and GPU allocations based on different service requirements and pricing models. The system enables flexible configuration of resource shares (e.g., 50:50, 40:60, 30:70) allowing customers to adjust the mix of CPU and GPU resources according to their specific performance needs and budget constraints, thereby optimizing the balance between processing performance and cost.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent implements partial action by allowing selective allocation of GPU shares based on actual processing needs rather than providing full GPU resources to all virtual machines. The system enables partial GPU allocation (e.g., 20-40% GPU shares) for workloads that don't require intensive graphics processing, reducing overall GPU resource consumption and associated costs while maintaining sufficient performance for the allocated tasks.

Inventive Principle:
Principle #16Partial or excessive action

3Quantity of substance

If the number of GPU shares is made greater than CPU shares, then the average cost of provisioning GPU resources is reduced, but the allocation becomes more complex and requires dynamic management

Engineering Contradiction:
Improveaverage cost of GPU provisioningVSAvoidshare allocation management
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the total processing resources into distinct CPU shares and GPU shares that can be independently allocated and managed. The system segments the resource pool into customizable share units, allowing flexible combination of CPU and GPU shares in various ratios for different virtual machines, which simplifies the management of asymmetric allocation by breaking down the complex provisioning process into manageable segments.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements universality by creating a unified share-based allocation framework that can handle both CPU and GPU resources using the same conceptual model. The system uses universal share units that can be combined in different proportions to satisfy diverse resource requirements, making the allocation mechanism versatile and adaptable to various workloads without requiring separate complex management systems for each resource type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250147790A1Methods, Systems and Computer Program Products for Optimized Virtualization of Processing Units for Cloud Computing Based Services
Publication Date: 2025.05.08 NOW GG INC
  • US20250147790A1 patent drawing
  • US20250147790A1 patent drawing
  • US20250147790A1 patent drawing

AI summary

The invention relates to on-demand cloud computing. In particular, the present invention provides methods, systems and computer program products for optimized virtualization of processing units for implementing cloud computing based services. In an embodiment, the invention includes (i) receiving a request for instantiating at least one cloud computing service, (ii) assigning for execution of the at least one cloud computing service a first set of central processing unit (CPU) shares, and a first set of graphic processing unit (GPU) shares, and (iii) executing a set of processes corresponding to the at least one cloud computing service using the first set of CPU shares and the first set of GPU shares, wherein the GPU shares and the CPU shares have been asymmetrically generated.