Modular HPC Resource Allocation with Dynamic Accelerator Assignment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current HPC architectures using off-the-shelf general purpose processors face limitations in performance, energy efficiency, and scalability, particularly in heterogeneous systems where the static assignment of accelerators to cluster nodes is inflexible and inefficient, leading to increased costs and potential component failures.
Innovation Solution
A modular computing system with a modular computing abstraction layer (MCAL) that dynamically assigns resources across different modules, including cluster, booster, storage, and other specialized modules, using node managers to manage and communicate resources efficiently, ensuring high resiliency and adaptability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Power
If the number of cluster nodes and processors is increased to satisfy computation power demand, then computation power is improved, but energy consumption increases and cost increases
Solution Approach 1:
The patent changes the architectural parameters by transitioning from traditional CPU-based clusters to GPU-accelerated nodes, fundamentally altering the computation-per-energy-ratio parameter. This allows achieving higher computation power with improved energy efficiency by utilizing the parallel processing capabilities of GPUs rather than increasing CPU count
Solution Approach 2:
The patent employs multiple GPU accelerators across cluster nodes, creating redundant copies of acceleration capability. This allows the system to achieve high computation power through parallel processing across multiple GPU units rather than relying on a single high-power processor, improving both scalability and energy efficiency
2Use of energy by moving object
If accelerators are attached to each cluster node to improve energy efficiency, then energy efficiency is improved, but the static assignment ratio between processor and accelerator reduces adaptability
Solution Approach 1:
The patent implements dynamic resource allocation where GPU accelerators can be dynamically assigned to different computational tasks and cluster nodes based on workload demands. The system allows accelerators to be moved, replicated, or reassigned during runtime, transforming the static accelerator-to-processor ratio into a dynamic configuration that adapts to varying computational requirements
Solution Approach 2:
The patent creates a universal accelerator resource pool that can serve multiple cluster nodes and different application types. The GPU accelerators are designed to handle diverse computational workloads through programmable architectures, allowing the same hardware resource to fulfill multiple functions across different scientific and engineering applications
3Productivity
If heterogeneous computer systems with different accelerators are used to improve scalability, then scalability is improved, but system complexity increases
Solution Approach 1:
The patent introduces a heterogeneous resource manager as an intermediary layer between the diverse accelerator resources and the applications. This manager abstracts the complexity of different accelerator types (GPUs, FPGAs, ASICs) by providing unified resource allocation, task scheduling, and performance optimization, allowing scalability without proportionally increasing system complexity
4Ease of operation
If off-the-shelf general purpose processors are used to ease manufacturing and programming, then ease of operation is improved, but energy efficiency deteriorates and computation power per cost decreases
Solution Approach 1:
The patent segments the computing system into two distinct functional parts: general-purpose processors for control and coordination tasks, and specialized GPU accelerators for compute-intensive workloads. This segmentation allows each component to operate in its optimal efficiency zone, with CPUs handling overhead and GPUs handling parallel computation, achieving both energy efficiency and maintained programmability through standard interfaces
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The present invention relates to the technical field of high performance computing (HPC). In particular, the invention relates to a heterogeneous computing system, particularly a computing system including different modules, which can freely be assigned to jointly process a computation tasks. A control entity, referred to as module computing abstraction layer (MCAL), is provided which allows dynamic assignment of various resources provided by the different modules. Due to its flexibility in adjusting to varying demands in computing, the present invention is also applicable as an underlying system for providing cloud computing services, which provides shared computer processing resources and data to computers and other devices on demand, mostly via the Internet.