Dynamic Hybrid Computing Environment Offloading
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Expensive, non-commodity computing resources such as GPUs are underutilized due to the inefficiency of current technologies in managing and allocating resources for computationally intensive tasks like deep learning, leading to high overhead costs and underutilization.
Innovation Solution
A dynamic hybrid computing environment is established, where virtual machines on commodity hardware clusters can offload tasks to high-performance clusters equipped with accelerators only when needed, utilizing a virtual network infrastructure for secure connectivity and data sharing through a distributed file system, allowing for efficient allocation and deallocation of resources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Power
If expensive non-commodity machines with accelerators are used for computationally intensive tasks, then computing performance is improved, but resource utilization deteriorates due to underutilization when special-purpose resources are not needed
Solution Approach 1:
The patent creates a hybrid computing environment where commodity hardware clusters can perform both general-purpose computing tasks and, when needed, offload to accelerators for specialized computationally intensive workloads. This multi-functionality allows the same infrastructure to serve multiple purposes, improving resource utilization while maintaining high performance when required
Solution Approach 2:
The system dynamically allocates and deallocates accelerator resources based on workload requirements. Virtual machines can be migrated between commodity and accelerator hardware, and resources are provisioned on-demand rather than statically assigned, ensuring that expensive resources are utilized only when necessary while maintaining computing performance for intensive tasks
2Quantity of substance
If commodity hardware is used for data preparation and model evaluation, then cost is reduced, but computing speed deteriorates for intensive tasks
Solution Approach 1:
The patent segments the computing workload into different phases: data preparation and model evaluation run on cost-effective commodity hardware, while only the computationally intensive DNN training phase utilizes accelerator resources. This segmentation allows the system to minimize costs for tasks that don't require high performance while allocating expensive resources only where they provide necessary computing speed
Solution Approach 2:
A virtualization layer acts as an intermediary between commodity hardware and accelerators, enabling seamless workload migration and resource orchestration. This mediator allows the system to automatically transition tasks between hardware types based on performance requirements, maintaining cost-effectiveness for standard tasks while enabling high-speed computation when needed
3Loss of time
If accelerators are deployed for deep learning training, then training time is reduced, but hardware cost increases due to specialized equipment requirements
Solution Approach 1:
The system applies accelerator resources partially and selectively only to the specific phase of deep learning training that benefits most from high-performance computation, rather than deploying accelerators for all computing tasks. This partial action approach reduces hardware costs by using expensive equipment only when the time-saving benefit justifies the expense
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Various embodiments herein each include at least one of systems, methods, and software for instantiating, executing, and operating dynamic hybrid computing environments, such as in cloud computing. Some such embodiments include allocating computing resources of a first server cluster to instantiate a first cluster and to establish a computing session. This embodiment may then initiate execution of a program within the first cluster that offloads at least one computing task to a second cluster, when the second cluster is instantiated, to leverage high-computing speed performance capabilities of the second cluster with regard to certain computing operations. Upon completion of program execution, the second cluster is then deallocated.